Friday, 27 March 2020

Technology Driven Insurance Data Analytics

The nature of the Insurance industry being data-centric, insurers abide by the policy of keeping data as a treasure for their respective growth. This data can only be turned into a gold mine of Insurance Data Analytics by in-depth analysis of other engagement areas of the customer and the insurer depending upon the kind of insurance one prefers.For example, when a car insurance of a customer is considered, every data from the social media profile of the customer to the kind of accident-prone zone the customer resides in is important and can help with better insights in the analytical prediction of risk management and even claim management for that matter. Humans have been exploring the lengths and breadths of data all this while but as the technological disruption of the world has taken away core complex tasks of humans, the insurance industry has started relying on the new-gen technology to deep dive into data – Big data and unstructured data – to pull out the best insights useful for them to help in various insurance business processes.The role of technology for extracting and in some cases, even implementing data analytical decision making is becoming a requisite for an insurer to grow. ...


Read More on Datafloq

Telcos set up war rooms to monitor networks during Covid-19 crisis

The three private telcos have obtained approvals from the DoT to ensure critical workers manning their NoCs can freely travel amid the pandemic-induced countrywide lockdowns

White House, Hospitals, Private Companies Exploring AI to Fight Coronavirus

By AI Trends Staff

The White House has issued a “call to action” to AI researchers to help fight the coronavirus spread; hospitals are pursuing AI to help manage the outbreak, and a number of companies are engaged in research to apply AI to the fight. Here is an update.

The White House recently announced an open data set with scientific literature on the novel coronavirus, called the COVID-19 Open Research Dataset or CORD-19, according to an account in TechCrunch. New data to be added to the centralized hub will be machine-readable.

US CTO Michael Kratsios called the new data set the “most extensive collection of machine-readable coronavirus literature to date,” in a press conference. He characterized the project as a “call to action: for the AI community. As a guide for researchers, the National Academies of Sciences, Engineering, and Medicine collaborated with the World Health Organization to come up with “high priority” questions about the coronavirus related to genetics, incubation, treatment, symptoms and prevention.

Michael Kratsios, US CTO

A number of organizations are participating in the partnership, including Chan Zuckerberg Initiative, Microsoft Research, the Allen Institute for Artificial Intelligence, the National Institutes of Health’s National Library of Medicine, Georgetown University’s Center for Security and Emerging Technology, Cold Spring Harbor Laboratory and the Kaggle AI platform, owned by Google.

The database is to include some 30,000 scientific articles about viruses in the coronavirus group. The database will include pre-publication research from resources including medRxiv and bioRxiv, open access archives for health sciences and biology research.

“Sharing vital information across scientific and medical communities is key to accelerating our ability to respond to the coronavirus pandemic,” stated Chan Zuckerberg Initiative Head of Science Cori Bargmann of the project.

Cori Bargmann, Head of Science, Chan Zuckerberg Initiative

Hospital Turns to AI to Help Monitor Patients, Visitors

Large healthcare systems are turning to AI to help monitor patients and regulate the flow of visitors as they work to contain the spread of the coronavirus. Tampa General Hospital in Florida recently installed a new AI system aimed at detecting visitors running a fever with a facial scan, according to an account in WSJPro. The hospital, which has more than a half-million visitors each year, recently admitted several patients with Covid-19 and is bracing for an influx.

The hospital recently installed an AI-powered screening system developed by Care.ai of Orlando, that uses devices embedded with cameras at the hospital’s six visitor entrances, to assess a person’s health by analyzing attributes of the face such as sweating and discoloration and by conducting a thermal scan.

The goal is to block those who have a fever from coming into the hospital, stated John Couris, president and CEO of Tampa General Hospital. Visitors are also asked about international travel and contact with people infected with Covid-19. The hospital hopes to reduce normal foot traffic by 75% and that it “keeps people that don’t really need to be in the hospital, out of the hospital,” Couris stated.

Private Industry Applying AI to Help Discover Coronavirus-Fighting Drugs

In private industry, drug discovery companies are putting AI technology to work to predict which existing drugs, or brand-new drugs could treat the virus. Five of them were profiled in a recent edition of  IEEE Spectrum.

The hope is to speed up a process that typically takes a decade to move from idea to market, with failure rates over 90% and a price tag of between $2 and $3 billion. “We can substantially accelerate this process using AI and make it much cheaper, faster, and more likely to succeed,” stated Alex Zhavoronkov, CEO of Insilico Medicine, an AI company focused on drug discovery.

Alex Zhavoronkov, CEO of Insilico Medicine

Drug development typically takes at least a decade to move from idea to market, with failure rates of over 90% and a price tag between $2 and $3 billion. “We can substantially accelerate this process using AI and make it much cheaper, faster, and more likely to succeed,” says Alex Zhavoronkov, CEO of Insilico Medicine, an AI company focused on drug discovery.

Based in Hong Kong, Insilico used an AI-based drug discovery platform to generate tens of thousands of novel molecules with the potential to bind a specific SARS-COV-2 protein and block the ability of the virus to replicate. A deep learning filtering system helped to narrow down the list.

“We published the original 100 molecules after a 4-day AI sprint,” stated Dr. Zhavoronkov. The testing was interrupted when 20 of the firm’s contracted chemists were quarantined in Wuhan, China, where the virus is believed to have originated. Still, the company has synthesized two of seven candidate molecules, which it plans to test further in coming weeks. The company has also licensed its platform to two large pharmaceutical companies.

British startup Benevolent AI recently identified approved drugs that might block the viral replication process of SARS-CoV-2. The company used a large repository of medical information extracted from scientific literature using machine learning, to identify six compounds that block a cellular pathway that appears to allow the virus into cells to make more virus particles.

One of the six is baricitinib, a pill approved to treat rheumatoid arthritis, which appears to be the best of the group for safety and efficacy against SARS-CoV-2. Benevolent’s co-founder, Ivan Griffin, stated that the company has reached out to drug manufacturers who make the drug about testing it as a potential treatment. Currently, ruxolitinib, a drug that works by a similar mechanism, is in clinical trials for effectiveness against Covid-19.

Study Finds CT Scans Can Detect Virus

An AI-powered CT image analysis solution designed to detect Covid-19 and quantify the disease burden in affected patients, has been validated by a research study to produce accurate results, according to RADLogics, which made the announcement.

A CT, or computed tomography scan is a medical imaging procedure that uses many X-ray measurements taken from different angles to produce cross-sectional images of a scanned object. An MRI, or magnetic resonance imaging, uses magnetic fields and radio frequency pulses to produce detailed pictures of organs. CT scans use radiation; MRIs do not.

The study, led by Professor Hayit Greenspan from Tel Aviv University, used an AI-powered image analysis platform for radiologists offered by RADLogics, a medical imaging analysis company with offices in Tel Aviv and Boston. Prof. Greenspan is Chief Scientist and a co-founder of the company, which started in 2010.

Prof. Hayit Greenspan, Tel Aviv University

The research team also included Dr. Eliot Siegel of the University of Maryland School of Medicine in Baltimore, MD; and Dr. Adam Bernheim of the Icahn School of Medicine at Mount Sinai in New York, NY.

The team developed a CT image analysis algorithm and trained it on multiple international datasets. They were able to differentiate 157 patients with and without Covid-19 with high accuracy. While not recommended as a first-line test, the CT test has been shown to be effective in detecting Covid-19 and following up with patients. The image analysis is also said to be effective in suggesting a “Corona Score” which measures the percentage of lung volume infected by the disease.

Moshe Becker, CEO and also a co-founder of RADLogics, stated in a press release, “As the novel coronavirus continues to rapidly spread around the world, healthcare systems and providers may become overwhelmed with symptomatic patients that require testing, imaging, and treatment. In an effort to help alleviate this burden on the world’s healthcare providers – and to support improved patient outcomes – we dedicated our resources toward successfully modifying and adapting our existing AI models to develop this solution specifically for Covid-19 detection and quantification. To date, we have deployed our solution in China, Russia and Italy, and we are rapidly scaling in other countries in response to the strong demand.”

Results of this study are available on arXiv.org. The study has been submitted to the Radiology Society of North America (RSNA) for review and potential publication in Radiology: Artificial Intelligence.

Read the source articles in TechCrunch, WSJPro and IEEE Spectrum; get more information at RADLogics.

Executive Interview: Beena Ammanath Boosts Women and Diversity in Tech as AI Expands

Putting Guardrails in Place for Women and Diversity in Tech as AI Is Infused Into Everything in the Organization

Beena Ammanath is the Founder and CEO of Humans For AI, a nonprofit organization focused on increasing diversity in tech leveraging AI. She is a recognized lead and industry expert who has driven pioneering technology changes in the use of AI, Data and Analytics for several market-leading companies. She has worked as a mentor to help women and minorities enter the new economy. She started Humans for AI in 2017 to help make AI understandable to the non-tech community.

Beena is also an Industrial Board Member of Cal Poly University, where she brings the industry perspective to influence curriculum engineers. She previously worked as CTO-AI at Hewlett-Packard Enterprise. She recently spent a few minutes talking with AI Trends Editor John P. Desmond.

Beena Ammanath, Founder and CEO of Humans For AI

AI Trends: Thank you, Beena, for being here with us today. I know you have been involved with data science and AI for many years in executive management positions for several major companies. Given your range of experience, what are the top trends around AI in business today that concern you or that you see having the most impact?

Beena Ammanath: Before it was called AI, there were statisticians and people who worked on predicting trends and looking at data, looking at managing big data and really driving insights. The term AI itself has existed since 1956. What we are  seeing now is its broad application and influence in organizations across the world.

Today AI is pretty much in every industry and every sector. Everywhere you look there are AI use cases. Some industries are more mature, more ahead with their AI use cases than the others. They tend to be the ones where there is more data easily available. In my career, I have worked in different industries, different domains, whether it’s financial trading, IOT, manufacturing, field services and telecom. It’s given me a broad perspective into how the data analytics and AI space has evolved.

We see many AI use cases in the financial sector and in the big tech social media sector. AI is not as prevalent yet in legacy  sectors like industrial manufacturing, Having to look at historical data becomes a challenge when you haven’t embedded the sensors to start with.

I have also seen a shift in the focus of AI. The focus up to two or three years ago was very much around value creation, doing POCs [proof of concept projects] and scaling up those POCs—“How do you actually drive value from the data?”—very focused around value and insights.

Now I see more and more urgency happening about the ethics part of AI, about what could potentially go wrong, and how do you put the guardrails in place? I have spoken about ethics for a few years now. It’s something that I worry about; I passionately care about making sure that what we build is the best thing for humankind. And I am so heartened to see more and more companies actually bringing up ethics earlier on in the conversation compared to say, two or three years ago.

You were the founder of Humans for AI in 2017. Can you say why you founded the group and how’s the work going there?

I have built a few data and AI teams and it tends to be very monochromatic. We all know there is a problem about lack of women in tech and diversity in tech. And I see it even more starkly reflected in the data and AI teams.

I worry about it, because we all have heard about how AI can be biased based on who builds it. Our biases get embedded into the AI system that we build. And the only solution to that is to bring more diversity to the table. And when I say diversity, I don’t mean just the gender part of it, but bringing in people from different backgrounds. Whether it’s a different economic background, professional background, educational background, people who think differently than you. That’s very important.

So diversity from every aspect is very important for us to build robust AI products and solutions. And having built these teams, I’ve noticed it’s really hard to find women or people of color or people from different education backgrounds to be part of the AI teams. And I see an opportunity for us to actually fix the problem. I also think that if we don’t fix the problem, then AI just won’t reach its full potential.

 

Humans for AI was founded with a single mission of increasing diversity in AI. Women tend to be the largest minority group, but it’s also about bringing in more people of color, people with different perspectives to the AI table—make all humans part of the AI narrative. So the way we are approaching it is by providing very focused AI literacy programs to diverse and minority groups.

We know that an AI team is not just the folks who have the PhDs in machine learning. Even though that’s what is communicated, I know they need to be surrounded by very good software engineers, very good designers and testers and project managers. So the way Humans for AI is approaching it, is by bringing in diversity in all the other roles that surround the data scientists.

We have a few programs as part of this initiative. (1) We have set up a foundation along with UC Berkeley called the Alliance for Inclusive AI; the foundation will be giving scholarships to women and minorities who want to study AI at UC Berkeley.

(2) We are also launching a virtual conference in May, which is going to be about teaching AI by profession. We will have a series around AI for specific professions: AI for nurses, AI for physical therapists, AI for elementary school teachers, for example. (3) We have Humans for AI ambassadors all over the world, who organize local meetups to help drive that mission. We are also partnering with other nonprofits who are more focused around teaching AI and coding to women and girls and minorities. So if anybody wants to study coding, we send them to our partner organization. All of this is anchored on a virtual community bring together the AI experts and novices on one virtual platform to interact and grow AI together.

It’s quite exciting. We are a bit behind the curve on getting more diversity into AI, but  this is my attempt to move the needle on it.

Very good. On the topic of how far along we are in the practice of AI, how mature is the industry today?

I have been on the board of several startups and I get invited to speak at board meetings as an expert on AI. I have seen a lot of interest. The trend has changed from what we can do with AI, to what we should be thinking about holistically to scale AI, to make sure that we are reaching the full potential with AI in our organization.

That includes the ethics piece as I mentioned earlier, but also, how you infuse AI to drive more value within the organization. It’s no longer just about being able to build better products, but how we can  use AI within every  functions in an organization, like finance or HR or legal. So being able to infuse AI across the organization, is a much more discussed topic today, compared to a few years ago. I also see many companies setting up separate innovation groups to look at not only just AI, but at other technologies like AR, VR and blockchain. AI is certainly a big part of the whole digital landscape, but looking at how the whole evolution impacts business is a much more talked about topic today.

I see you are an advisor to the California Polytechnic State University, College of Engineering. How’s that going?

I have been on the Cal-Poly industrial advisory board for quite a few years now. What we teach the next generation of engineers has to be very much relevant to what we are seeing in the industry.

The best way to shape the thinking of future generations, is by getting more actively involved in how the curriculum is shaped, how it is delivered, and bringing in industry engagement. I have definitely learned a lot by engaging with Cal-Poly, and I hope my input is helping to shape the future of tech curriculum.

Sounds good. How do you believe the workforce is responding to the challenge of learning about AI, while not getting overrun by automation?

We definitely see a lot of hype around AI. We are used to seeing headline articles about AI taking away everybody’s job and jobs being eliminated completely. The reality is, that is hype. There is a grain of truth, in that certain jobs will get eliminated, but AI is also creating a ton of new jobs. And you don’t hear so much about that. You hear a lot about jobs getting eliminated.

I’m a history buff. When I look back, the jobs that existed 100 years ago look completely different from the jobs that exist today. Not only from a title perspective or the role perspective, but even how the job was done was very different 100 years ago compared to today. Or even 50 years ago compared to today.

So there will be job elimination, and there will be job creation as well. I have definitely seen a bigger need for more people with technical chops. Think about all the sensors getting embedded into all the physical devices we have. Not only our phones, but our cars, our microwaves and our refrigerators – more and more compared to even a few years ago.

We need more hardware engineers. If you search, you will find job descriptions for sensor cleaners for the automotive industry. That’s a new job. If you think about the traditional AI technology teams, more people are needed for the data curation and the validation of the AI results. Those jobs didn’t exist 20 years ago. More roles will get created. As automation takes over some of the more mundane and boring tasks, we will see many new jobs getting created as part of the AI journey we are on.

Do you have any advice for young people or mid-career people interested in pursuing AI? Where should a college undergrad start? What would you recommend for an early or mid-career professional?

If you are already on a career track within technology, it is absolutely important that you know about AI and you have basic AI literacy chops. It’s like any other technology, but I do think basic AI literacy is important for any professional. Anybody who is part of the workforce today needs basic AI literacy, just because AI is so infused into everything that we do. No matter what your job is, you’re going to be using AI in some form or the other. So I think basic AI literacy is super important. Lots of curriculum is out there and organizations like Humans for AI are really focusing on making AI literacy more and more accessible.

If you are a student who wants to go down the AI career path, there is the traditional pure data scientist part of studying machine learning and AI, going very deep into that. But there’s also, as I was saying earlier, a lot of demand for people who design the AI solutions, who make them more accessible – UX designers for AI and testers in QA for AI for example. Many paths can enable you to work in AI and the roles can be different. Even as we speak, new roles and titles are getting created in the AI job market.

Learn more at Humans for AI.

Clinical Data Sharing for AI: Proposed Framework Could Rouse Debate

By Deb Borfitz, Senior Science Writer

A group of doctors from Stanford University has proposed a framework for sharing clinical data for artificial intelligence (AI) that could set off a firestorm of debate about who truly owns medical data, ethical obligations to share it, and how to properly police researchers who use it. On the other hand, the envisioned approach has parallels to the open science tactics currently being uniformly deployed to battle the COVID-19 pandemic.

The framework’s central premise is that clinical data should be treated as a public good when it is used for secondary purposes such as research or the development of AI algorithms, as detailed in a special report (doi: 10.1148/radiol.2020192536) published recently in Radiology. That means broadening access to aggregated, de-identified clinical data, forbidding its sale and holding everyone who interacts with it accountable for protecting patient privacy, explains study lead author David B. Larson, M.D., M.B.A., vice chair of clinical operations for the radiology department at Stanford University School of Medicine.

Dr. David B. Larson, Vice Chair, Clinical Operations, Department of Radiology, Stanford University School of Medicine

Although the framework published in a journal specific to radiology, and three of its authors are radiologists, the structure is “universally applicable to other types of medical data as well,” says Larson.

Disputes over clinical data sharing generally involve those who believe patients own the data and those who think institutions do. But Larson and his colleagues advocate a third approach, saying nobody truly owns the data in the traditional sense once it has served its primary, patient care purpose—whether that happens immediately or in another 20 years.

“When data are aggregated and deidentified, and insights get extracted… that is a separate activity,” he says. “That’s not what it was designed to do initially but it is a fortuitous secondary use.”

The doctors further argue that patients, provider organizations, and algorithm developers all have ethical obligations to help ensure that these clinical observations are used to benefit future patients. “We now have to develop some thinking around how we are going to address the appropriateness of that use.”

The current COVID-19 pandemic is a “pretty clear,” if unfortunate, illustration of what a world might look like with an ethical framework in place so patients can trust that their data won’t be used inappropriately and researchers aren’t stymied in their efforts to use it for clinically beneficial purposes, says Larson. While the framework was written in the spirit of open science, he adds, it also recognizes that algorithms derived from clinical data may have intellectual property associated with it—which should not preclude other people from having access to the same raw materials.

Irreconcilable Differences

Larson and his Stanford colleagues felt pressed to develop an ethical foundation for clinical data sharing in the absence of any formal guidance for holding themselves accountable, Larson says. “Once we did that it seemed reasonable to share it with others.”

The framework is designed to overcome the irreconcilable positions of those who say either patients or providers “own” the data—meaning, who has permission to access it and rights to profit from it—which is “preventing us from moving forward in a reasonable way,” says Larson. The views of neither camp can be fully justified from an ethical standpoint. “We subscribe to an ethics framework (doi: 10.1002/hast.134) developed by Ruth Faden and others that we all should be contributing to the common good.”

Since people have been benefitting from the research and improvement efforts of health systems for centuries, he reasons, they should not be able to withhold the use of data in ways that benefit others in the future. “We think we’re providing an avenue that addresses the major concerns on either side.”

The framework will hopefully lay the groundwork for future, national-level changes to the way both data and organizations are structured, in and outside the U.S., says Larson. It was only relatively recently that discrete data in electronic health records were even available for use, and the tools to process and learn from that data have been on the scene for even less time.

“Patients generally don’t withhold data from their care provider when it is being used on their behalf … [because] a relationship of trust has been established,” Larson says. Yet when people think clinical data is being used by AI developers and other non-provider entities, they’re immediately suspicious.

“We’re pushing back on that and saying if those entities can’t currently be trusted then let’s create the ground rules so they can be… and if they’re not willing to participate in that environment then they shouldn’t have access to the data. Let’s increase the inherent trust in the system by holding those who have access to the data accountable to be good stewards,” as providers are already doing.

“We can hold ourselves accountable and the broader community more accountable for being good data stewards, which hasn’t really happened up until now and we think it should,” says Larson. As envisioned, the entity releasing the data would be responsible for ensuring that it is going to a trusted partner and being used as contractually specified.

It’s “reasonable” for providers who maintain and process clinical data to charge outside entities an access fee, but it should not be excessive, Larson continues. But under no circumstances should they strike an exclusive agreement with one entity that precludes the same access by others.

Regulations will also need to be written to further ensure entities use the data appropriately and for beneficial purposes, Larson says. This would include penalties for using data they receive that is accidentally identified or using technology that allows them to identify individuals from the data.

As Larson points out, the framework refers to “wide” rather than public release of data in the belief that data should be used only by those who identify themselves and agree to be held accountable for its appropriate use.

Early concepts that contributed to the proposed ethical framework were presented at BOLD AIR (Bioethics, Law, and Data-sharing: AI in Radiology Summit), organized by the departments of radiology at Stanford and New York University Langone Medical Center last April, says Larson. The one-day event was co-sponsored by the American College of Radiology, Radiology Society of North America, Massachusetts General Hospital, Stanford Center for Artificial Intelligence in Medicine and Imaging, and the Center for Advanced Imaging Innovation and Research.

A follow-up meeting to further discuss the salient issues is now on hold due to COVID-19. If nothing else, Larson says, the proposed framework is in the literature to fuel thoughtful discussion and hopefully inform future regulation.

Outstanding Questions

The strongly worded announcement of intention came as welcome news to Megan Doerr, a genetic counselor and principal scientist with the open-science organization Sage Bionetworks. “The more people are able to use scientific data for scientific solving, the more scientific solutions we’re going to have to our problems,” she says.

Megan Doerr, Principal Scientist, Sage Bionetworks

“The devil is really in the details,” Doerr continues. While she applauds the idea of extending responsibility for the data to anyone who uses it, for example, it is uncertain how that might be practically accomplished.

Institutions have traditionally acted as proxy bonding agents for researchers, so there is “someone to fine or sanction” if there are ethical violations, Doerr says. Who would serve as the ethics watchdog once large data sets are opened to a larger, more diverse community of users? Maybe Lloyds of London, which bonds astronauts on space shuttle missions?

Another concern is how to appropriately protect the privacy and rights of people whose data are being shared, she says. “The more data that is available and can be cross-referenced, the quicker we realize ‘de-identified’ data is not [really] de-identified… We can’t promise anybody that their privacy is going to be protected and as scientists we need to be honest about this, and we are not. We contort in a million different ways to avoid this uncomfortable truth.”

What’s needed are legal protections so if the information is used in ways that are inconsistent with the agreed data use, money would flow to people who were impacted to mitigate the harm, says Doerr.

The proposed framework may also raise social justice and equity questions, she adds, since “brilliant but resource-limited scientists” may not be able to afford the entry cost of problem-solving.

Cost Concerns

The paper talks about a lot of important issues, offers a sound ethical framing for clinical data sharing and is well cited, says Doerr. But she views the proposed framework more as an “opening salvo” due to what it does not address—who might pay for the cost of compute and metadata harmonization, which are the two biggest barriers to more effective AI research.

Radiological datasets are massive, measured in petabytes, and therefore require a tremendous amount of computing power to host, she notes. “Nobody can download the data; it takes forever. “So it’s not like researchers are going to be downloading a local copy of the data to work with… and to create a hosting space for these data is a very expensive thing to do, which is one of the challenges of the All of Us Research Program [of the National Institutes of Health].”

Doerr chairs the researcher application subcommittee for the All of Us Research Program and sits on the resource access board for the All of Us dataset. David Magnus, Ph.D., a professor of medicine, biomedical ethics and pediatrics at Stanford, and one of the authors of the proposed clinical data framework, is vice-chair of the institutional review board for the program.

Doerr wonders: Will researchers have to pay into a system giving them access to a sandbox area where the clinical data are stored? If so, which institutions would be doing the primary data gathering?

Authors of the Radiology report make the point that aggreging clinical data from multiple institutions “may markedly enhance the value of the data,” Doerr says, “which is absolutely true, but… incredibly expensive and difficult. So, who is going to do that work and who is going to pay for it?”

It would be a “tremendous waste” to have individual AI developers be responsible for hosting their own data and doing their own metadata harmonization, says Doerr. More importantly, it could lead to inconsistent results. The datasets they’d be working with would be “effectively tuned into different keys and may not return compatible insights.”

This is a problem Sage Bionetworks has been toiling over for quite a while, Doerr says—as has study co-author Nigam H. Shah, MBBS, Ph.D., associate professor of medicine (biomedical informatics) at Stanford and assistant director of the Center for Biomedical Informatics Research.

The paper alludes to data stewardship and talks about federated learning, a system Sage Bionetworks used for its Digital Mammography DREAM Challenge that demonstrates both the efficacy and challenges of the approach, she says.

“We had hundreds of thousands of digital mammography images and recognized that they couldn’t be de-identified, so we had solvers send their machine learning models to us and we then ran them against the data on their behalf and returned the results to researchers. In this way, they never saw the actual data, but they could still fit their models to it.”

The Challenge was a costly undertaking supported by grants, Doerr says. Scientists at Sage spent an exhaustive number of hours manually harmonizing the datasets so that the data returned authentic, trustworthy results.

Very few people are good at metadata harmonization, she adds, because the cost of compute is so expensive it limits who gets to do AI research. “Honestly, that might be an OK thing right now. I don’t think communities have any idea about the individual and group harms that could be caused by bad AI research… [that] within medicine could really be a problem.”

In discussing their proposed framework, the authors say initiatives such as the All of Us program could serve as an example of “how to allow participation from any qualifying research and development organization following an established vetting process.” The reality is that the researcher application subcommittee is a 25-person team effort with an expected development timeline of two to three years, Doerr says.

On the cost of compute question, Larson says the cost should be borne by the party that accrues the benefit. “If it turns out that the value is mainly to the public, then maybe this should be financed like other research through public and private entities.” Alternatively, commercial entities might pick up the tab if they are profiting from the intellectual property that they derive from the data. “I think there are a number of potential finance models that do not require selling of the data.”

Figuring out who will do the data harmonization work will be an iterative process, he adds. “I think it would be unwise and almost certainly untenable to try to impose a single standard right now because I think there will probably be other purposes that will drive standardization over time more quickly, such as  allowing patients to move from one healthcare system to another. There will likely also be other processes to help reconcile multiple standards.”

Pushback Expected

Some medical institutions might try to claim indefinite clinical data ownership by arguing that the information serves important care purposes for patients’ lifetime—or longer, given the implications for family members, says Doerr. But from her perspective, everyone owns a copy of the data; they just need it for different purposes.

“Data are not a traditional commodity,” she says. “I can have a copy, the hospital can have a copy and there can be a copy out in the public domain and all of it still retains its usefulness.” Ultimately, more users only magnifies the value of the data.

“The whole concept of ownership is somewhat flawed,” Doerr says. “As an individual, I have a right to the data that is generated about my body from my body, and in consenting for medical care I give my providers the right to that data, too. I am paying them for the service of interpretation of those data.” The pool of data that gets generated by this service over time “should flow into the public domain and be used to the benefit of the public good.”

Doerr says she feels strongly that everyone has an ethical obligation to ensure that happens, a concept that is more intuitive in nations with a single payer system such as Canada and the United Kingdom. “Our system because of its byzantine structure makes it a little less obvious, but our ethical obligation remains constant.”

This altruistic contribution of data for societal benefit is precisely what is happening in the wake of the coronavirus, “because we have an emergency… [and] we know data sharing can help,” Doerr continues. “One could argue that breast cancer that kills way more people every year might be a similar emergency and there are many other conditions that are equally deadly if not more so.”

If an open science approach can be embraced on a worldwide scale for the COVID-19 outbreak—including a worldwide commitment to make research and data freely available—”why can’t we do it for anything else?” she asks.

Notification to use clinical data should be required, and not just on an individual level, since privacy protection cannot be guaranteed, she contends. “If we allow folks to opt out of the system the data will become even more biased than it already is, and artificial intelligence and machine learning techniques only serve to amplify that bias.”

What’s needed is a “collective conversation” resulting in society-wide consent to this new paradigm acknowledging the value and accepting the risk related to the secondary use of clinical data, she continues. Large-scale genomic data sharing has familial implications that no amount of goodwill can wholly safeguard via the standard informed consent process, especially when identical twins don’t share the same sentiments about dating sharing.

“That’s why lawmaking is going to need to be the central component of this,” Doerr says. “Laws are one of the ways we exert our collective will toward a given end.”

While Doerr enthusiastically supports the level of clinical data sharing envisioned by Stanford radiologists, she notes that a multitude of financial incentives are working against their ideas—notably, the multi-trillion-dollar data brokerage business. Some pushback can be expected from large medical institutions in the U.S., including the not-for-profit Mayo Clinic and Cleveland Clinic that generate a sizable amount of revenue selling data to Google. As was widely reported last November, and referenced in the Radiology report, Google also made a widely debated deal with the 150-hospital Ascension health system to further its AI agenda.

Learn more in the special report (doi: 10.1148/radiol.2020192536) published recently in Radiology.

Quantum Computing with AI Seen Helping to Advance IoT

By AI Trends Staff

While still in a development stage, quantum computing is starting to advance into a new era, where it is poised to accelerate AI and the Internet of Things.

Quantum computing is seen as helping us to address some of the biggest and most complex challenges we face as humans, suggested author Chuck Brooks in a recent account in Forbes. Brooks is a thought leader for cybersecurity and emerging technologies, and a chair in the Quantum Security Alliance, formed to allow academia, industry, researchers and government to collaborate.

Quantum computing comes along at the right time for the Internet of Things, the idea that nearly every electronic device is addressable on the internet. The number of connective devices is expanding rapidly; an estimate from Business Insider Intelligence is that 40 billion IoT devices will be in place globally by 2023, including sensors, data, machines, people and the interactions between them.

The challenge is how to monitor and ensure quality services from all these devices. “Responsiveness, scalability, processes, and efficiency are needed to best service any new technology or capability. Especially across trillions of sensors,” Brooks wrote.

Quantum technology can also be helpful in addressing network latency, interoperability, AI, real-time analytics, predictive analytics, increased storage and data memory, secure cloud computing and the emerging 5G telecommunications infrastructure.

“As quantum computing and IoT merge, there will also be an evolving new ecosystem of policy issues,” he suggested. These include, ethics, interoperability protocols, cybersecurity, privacy/surveillance, complex autonomous systems, and best commercial practices.

Quantum Theory

The foundation of classical computers is transistors, which process in a binary system where bits are marked with a 0 or a 1, on or off, and so on. Quantum theory suggests it is possible for two things to be in two places, each at the same time. This theory can be used within computing to create a more complex and powerful system, suggested a recent article in DisruptionHub written by Dr. Chloe Sharp, managing director of Snap Out market consultants.

Dr. Chloe Sharp, Managing Director, Snap Out

Quantum computers use “qubits,” which can be any proportion of 0 and 1 at the same time, so they can process more information more quickly. Use cases range from facial recognition and complex database searches, to machine learning and safer encryption.

While quantum computing is still in the R&D stage, it’s beginning to make its way into several markets. The author suggests that for quantum computing to be widely accepted, the user experience (UX) needs to be improved, especially around creating trust with the user.

“In our experience within emerging technologies, it’s easy for those creating products and services to forget that they are dealing with tech that is extremely complex. As such, it is absolutely crucial that usability testing is done on all quantum and IoT products to ensure that they are accessible and easy to use, and solve real-life user problems,” Dr. Sharp suggested.

A degree of uncertainty surrounds the security of quantum computing.

Interest in Quantum Key Distribution (QKD) is expected to spike in 2020, as interest in the security methods surrounding quantum computing gain interest, suggests a report from CB Insights cited in IDQ, a company focused on security for quantum computing. QKD is a security communication method that enables two parties to produce a shared, random secret key known only to them, which can be used to encrypt and decrypt messages.

AI will be made profoundly more powerful if backed by quantum computers, the report suggests as well.

Microsoft and Amazon recently announced they are entering the quantum computing market, and Google has laid claim to “quantum supremacy” as a result of its advances, further indications of the acceleration happening.

Read the source articles in Forbes, DisruptionHub and at IDQ.

The Debate About Electric Vehicles (EVs) and AI Autonomous Cars

By Lance Eliot, the AI Trends Insider

Electrical Vehicles (EVs) are talked about, they are praised, they get a lot of attention, and in some parts of the United States there is a near obsession with them (hint: California).

In spite of all the hype and press, the reality is that there are only around 1.1 million such cars in the U.S. and it represents a small fraction of the 250+ million cars in the country. That’s less than one-half of one percent of the total cars in circulation.

When I say this at various industry presentations, those with an EV are quick to yell at me as a traitor and get upset at my seemingly naysayer commentary.

Allow me to clarify that I am fully supportive of EVs and hope that a lot more will get sold. I’m a big cheerleader for EVs. All I’m trying to point out is that we have a long way to go before they become prevalent.

In terms of EVs, I am lumping together all variations in this herein discussion, for convenience’s sake. Generally, there are Plug-in EV’s (PEVs), consisting of Battery EV (BEVs) that are equipped to only run on batteries, and there are the Hybrid EVs (abbreviated as either HEVs or PHEVs), which use both a gas powered internal combustion engine and battery power.

Why have EV’s? One argument in favor of EV is that they are less polluting than conventional gas-powered cars. Thus, ecologically, the EV is better for the environment. We can all breathe a bit easier. Another argument is that the adoption of EV’s might aid in reducing the pace of climate change. That’s one that gets a lot of people in a tizzy since there are some that believe in climate change and some that do not.

Here’s a less controversial point, EV’s would reduce the dependence on oil and the production of gasoline. This would seem like a handy move since there are various predictions about how costly it is coming to become to get oil and make gasoline. Presumably it’s a limited resource, and we’re using it up. Also, it obviously tends to provide power to those that have it and not so much to those that don’t. Some say down with the cartels.

Another less discussed aspect in favor of EVs is that people like the quietness and the feel of driving an EV. I’m going to list that as an argument in favor of EVs, but I realize not everyone necessarily likes that aspect. There are some that love the sound of a conventional combustion engine and refuse to get an EV because it doesn’t have the same sound and fury. To each their own.

Of course, some EVs can simulate the sounds of a conventional car, or make other purposeful noises or sounds to serve as a warning or indicator that the EV is nearby.

The government right now is offering incentives to have people buy EVs, so from that perspective the government is considered somewhat supportive of EVs, which helps to promote them and keep the price lower than presumably what it might otherwise be.

Now, hold your breath, here are some of the stated negatives about EVs.

Some would say they are too expensive. Plus, if the government reduces the incentives to get one, it will be even more costly. Of course, the counter-argument is that we are still in the early days of EVs and presumably, eventually, the cost will come down.

The automakers are right now pretty much taking it on the chin to develop, make, and sell EVs. For example, news reports suggest that the Chevrolet Bolt has allegedly been often sold at a loss. The Nissan Leaf has been claimed to not be making a profit. I think it’s fair to say that right now all the auto makers that are into EVs are finding themselves faced with razor sharp margins and it’s quite a feat to find a profit in this, so far.

That being said, one could look at this as a wide-open market. It’s poised to explode, some argue. Currently in its infancy, one would expect that the adoption rate is low and the costs are high at the start of any new innovation adoption.

At some point, the popularity goes up and the price will be coming down. Most would claim that they can see on the horizon a mass market electric car that turns a nice profit. Sometimes you’ve got to invest in something at a loss, being patient before it turns around and hopefully becomes a true money maker.

The auto makers have to play the game since otherwise, if the market does become EV crazed, each automaker will need to have its own EV for consumers to buy. Imagine if the EV market booms and you are the only automaker that didn’t have the foresight and fortitude to put together an EV. That would be bad news for you, for your company, for your shareholders, etc.

You could add that having an EV is considered politically correct too. For some people, they enjoy bragging about their EV. It is considered stylish. Want a piece of the future, today? Get yourself an EV.

There are political analysts that worry about which country will get into EV first. China right now is going gangbusters over EV. Should we look at dollars or should we look at units sold?

There are some focusing on luxury EVs, notably having a much higher price tag than the everyday EV. There are arguments about which automaker is in the lead, depending upon your metric of using dollar sales volume versus number of units sold.

What does this have to do with AI self-driving driverless autonomous cars?

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. Doing so also makes us aware of the electrical power needs of an AI self-driving car. During my presentations at industry conferences, attendees often assume that all AI self-driving cars will unarguably be EVs. There is a bit of shock when I point out that this is not necessarily the case.

Here’s what people seem to say:

  • EVs will be the cause of AI self-driving cars, without which there won’t be AI self-driving cars.
  • AI self-driving cars will be the cause of EVs, without which there won’t be EVs.

Neither of those statements make much sense when you take them apart or unpack them.

Let’s tackle the notion that EVs will be the cause of AI self-driving cars (and, the added corollary that without EVs there won’t be AI self-driving cars).

It’s a kind of hyper claim that mishmashes things together.

We’ll begin with some fundamentals.

An AI self-driving car has lots of sensory devices, such as cameras, radar, sonic, LIDAR, and the rest.

These all require electrical power to run. An AI self-driving car has lots of computer processors and memory devices which are needed to run the AI part of things. These all require electric power to run.

It’s readily apparent that an AI self-driving car needs a lot of electrical power in order to work.

Not much debate on that.

For my article about LIDAR and other sensors see: https://aitrends.com/selfdrivingcars/lidar-secret-sauce-self-driving-cars/

For my framework about AI self-driving cars, see: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

Where is the AI part of the self-driving car going to get all of this needed electrical power?

Somehow, the car has to generate it.

If you use a conventional gas-powered car, you’ll likely need to outfit the car with additional electrical generation and power storage capabilities to meet the demand of the AI and its sensors.

This can be done.

You can argue that it raises the cost of the gas-powered car, which that’s likely true, and you can argue that it will take up space in the car, which is also likely true.

So, yes, a conventional gas-powered car might need to chew-up the trunk space to have added batteries and electrical elements, and overall the cost of the car is likely to go up.

All probably true.

The point is that it doesn’t preclude the use of a gas-powered car to be used for an AI self-driving car platform.

It just means that a gas-powered car is perhaps a less amenable choice.

There are some that are trying desperately to create kits that could turn a conventional gas-powered car into an AI self-driving car, which if this could be done would be a bonanza since you could sell the kit to presumably the 200+ million car owners in the US today.

Unfortunately, the kits are not likely to be viable, mainly because of the add-on needed to a conventional car, and not solely having to do with the electrical power constraints.

For my article about kits and AI self-driving cars, see: https://aitrends.com/selfdrivingcars/kits-and-ai-self-driving-cars/

So, let’s go ahead and reject the notion that AI self-driving cars are not possible without EVs.

That just doesn’t ring true.

The first part of the statement was that EVs will cause the advent of AI self-driving cars.

That’s only half true, I’d say.

I think it’s fair to say that even if EVs didn’t exist that we would all still be pouring our hearts into trying to create AI self-driving cars.

The electrical power aspect is for most AI developers an afterthought.

They aren’t worried about whether this machine learning system or that AI code is going to require a heftier processor that consumes more power.

Their assumption is that the power will be found. It’s up to those clever automotive engineers to get them the power needed.

For my look at AI self-driving cars as a moonshot, see: https://aitrends.com/selfdrivingcars/self-driving-car-mother-ai-projects-moonshot/

I remember when smart phones first came out.

The amount of available battery power was negligible. It seemed as though if you used your smartphone for the running of a game app for a few minutes and if you made a phone call or text, voilà you were out of power. Power consumption really wasn’t much concern per se.

Of course, consumers were irked and the makers of the smartphones woke up and realized that consumers would choose possibly one brand over another based on how long the battery lasted. This launched an intense interest in the battery makers and also how to optimize the OS for the lengthening of the battery life.

I’d wager the same is going to happen with AI self-driving cars.

At first, they’ll chew-up electrical power like they need the Hoover Dam or a nuclear reactor.

Once it’s been shown that AI is working and we believe in self-driving cars, the attention will shift toward using less power when possible and extending the batteries of the self-driving car as long as possible.

That’s though a second or third step in the evolution of things.

Thus, I don’t think you can say that the rise of EVs will “cause” the advent of AI self-driving cars.

The word “cause” is a pretty strong one.

If I stand next to you, and I shove you, and you fall to the ground, I’ll grant you that I “caused” you to fall down.

On the other hand, if I stand next to you, and you happen to fall down, and my standing next to you was a contributor (maybe you thought I was going to shove you and so you preempted it by dropping to the ground), I’d say that I was involved and there was some kind of correlation or relationship between the two aspects, but one wasn’t the cause for the other per se.

The advent of EVs is going to make the production of an AI self-driving car easier and hopefully less costly.

This is due to the aspect that the EV is already geared up to produce and store quantities of electrical power.

Therefore, the AI systems and sensors can tap into it. Or, if it is needed to boost the electrical capabilities for aiding the added AI components, it would seem a natural extension of what the EV already has for its design and construction.

You can also perhaps say that without the advent of EVs, it would likely make the advent of AI self-driving cars harder, more costly, and maybe even delay their rise.

Trying to graft onto a conventional gas-powered car the needed electrical storage and generation might distract from the AI side of things. The power engineers might say to the AI developers that they need to cut back on things. A conventional car might need to be redesigned and maybe even made to be bulkier or more awkward in shape and weight. Those are bad and unfortunate possibilities, but none of those would kill the desire to get to an AI self-driving car.

I’d dare say that there is such an intense desire to get to an AI self-driving car that if needed the auto makers would start over and make some kind of new car to accommodate it.

This doesn’t seem to be needed as yet.

So far, it appears that with pretty much a normal car, whether an EV or a gas-powered one, we’ll be able to make it into an AI self-driving car.

And, though the EV advocates will get angry at me, I’d claim there is as much if not even more calls for an AI self-driving car than there is for an EV.

Ouch, I said it.

People kind of buy into the EV advantages of the environment and all the rest, but it doesn’t seem to pique their interest. It’s a car that happens to use electricity. Nice. That’s a good thing.

On the other hand, if you tell someone that we can make a self-driving car, which will drive you whenever you want, and you don’t need to drive the car, it’s something that people say, yeh, I want one of those. It has the draw of allowing for a sense of freedom. It will provide mobility for the masses. It will change the nature of society. An EV is not going to do the same, sorry.

This brings us to the other statement that I had mentioned earlier, namely the claim that perhaps AI self-driving cars will bring about the advent of EVs, or as stated “cause” it to occur.

Kind of.

As I’ve already mentioned, it is not a precondition that an AI self-driving car can only happen if the car itself is an EV. There isn’t a requirement that the underlying car must be an EV. If it was the case that only an EV would suffice, I’d say that there would be even more strenuous efforts to get the world toward EV, as being sparked by the desire to get us to an AI self-driving car.

Will AI self-driving cars spur the advent of EV?

Yes, I think we can say that without much reservation or qualification.

The alignment of the need of the AI to have lots of electrical power with an underlying platform made for that purpose, the two will certainly fit like a glove. They will go hand-in-hand, as it were.

I’ll once again quibble with the notion that the AI self-driving car will cause EV adoption.

If the AI self-driving car makers opt to use EV as the underlying platform, and if the sales of those self-driving cars goes through the roof, due to the AI part of things, I’m not sure if we’d call that a cause-and-effect per se. Anyway, with the rising tide, all boats rise, as they say.

We probably should also consider the other side of the coins on this discussion about EVs and AI self-driving cars.

Suppose that we aren’t able to achieve a true AI self-driving car?

Meanwhile, suppose that the auto makers and tech firms have been using EVs as the primary platform for this AI self-driving car attempts. Could this hurt the advent of EVs?

Maybe people would get confused and inextricably consider the AI and the EV as one thing.

Thus, if the AI doesn’t cut the mustard, perhaps consumers would partially blame the EV side.

Out goes the baby with the bath water.

Suppose the EVs cannot come down in cost and remain relatively expensive.

If the AI self-driving car is tagged on top of the EV, it presumably now becomes even more expensive.

Would this price out the AI self-driving car for the masses?

Would people perceive the AI self-driving car as an elitist toy?

This could generate a backlash against AI self-driving cars.

Another consideration is the range of the EV.

If it is consuming electrical power to run the car and to run the AI portions, it might need to frequently stop at a charging station to recharge. If you add to this the notion that many believe most AI self-driving cars will be used non-stop 24×7, and that they will become a dominant ride sharing mechanism, the cost to stop and thus lose money while sitting there and charging, well, it is going to hurt. In spite of the downsides about gasoline, the length of time to fill-up is minimal and the distance drivable is high.

For my article about the nonstop desire for AI self-driving cars, see: https://aitrends.com/selfdrivingcars/non-stop-ai-self-driving-cars-truths-and-consequences/

The EV world is working on these aspects of reducing the charging time and maximizing the distance an EV can go.

Meanwhile, though, it’s another consideration about having an AI self-driving car that becomes intertwined with an EV platform.

Since the EV and the AI self-driving car might be perceived as joined at the hip, given that both are rising up at about the same time, whatever happens to one can spill over into the other.

Imagine an AI self-driving car that goes awry and kills someone.

Will consumers and the public realize that this is perhaps due to the AI and has nothing to do with the EV?

If the public cannot separate the two, it could cause a black-eye on EV, even though this would presumably be completely unfair and unwarranted.

Sometimes when I’m on a plane flying to give a speech, the person in the seat next to me will be curious when I say that I’m in the midst of developing AI self-driving cars.

They will often say something like they wish they too had an electric car, stating as such under the assumption that EV and AI self-driving cars are one and the same.

The odds are that they will be, and we can pretty much go along with the idea that AI self-driving cars are going to be EVs. I usually just smile and mumble that yes, EVs are cool.

For those of you doing automotive power engineering, as an AI developer, all I can say is “Scotty, we need more power” (in the immortal words of Captain Kirk).

I thank you in-advance for doing so.

Copyright 2020 Dr. Lance Eliot

This content is originally posted on AI Trends.

[Ed. Note: For reader’s interested in Dr. Eliot’s ongoing business analyses about the advent of self-driving cars, see his online Forbes column: https://forbes.com/sites/lanceeliot/]

Supercomputer Completes Massive Coronavirus Calculations

Scientists are fighting back against COVID-19 with the help of supercomputers. Using the Frontera supercomputer project, researchers are developing a coronavirus model to learn more about its effect on the body. This research could help scientists develop a treatment for the disease.With an accurate simulation, scientists could test how the virus reacts under various circumstances. The model would have to include how every atom interacts with each other to be accurate. That's a challenging prospect, considering there may be as many as 200 million atoms involved.Between March 12 and 13, researchers ran the initial simulations on Frontera. These simulations, which are just a fraction of what the final model will be, ran on more than 200,000 processing cores. This kind of processing power would be impossible without a computer like Frontera.The Frontera SupercomputerThe Frontera supercomputer, located at the University of Texas, is the 5th most powerful computer in the world. The National Science Foundation (NSF) granted the Texas Advanced Computing Center (TACC) $60 million to build Frontera in 2018. When researchers from the University of California began this coronavirus project, they immediately turned to Frontera.Even on a supercomputer like Frontera, these simulations are challenging. To account for the sheer ...


Read More on Datafloq

Thursday, 26 March 2020

New Methods for Improving Supply Chain Demand Forecasting

Organizations Are Rapidly Embracing Fine-Grained Demand Forecasting

Retailers and Consumer Goods manufacturers are increasingly seeking improvements to their supply chain management in order to reduce costs, free up working capital and create a foundation for omnichannel innovation. Changes in consumer purchasing behavior are placing new strains on the supply chain. Developing a better understanding of consumer demand via a demand forecast is considered a good starting point for most of these efforts as the demand for products and services drives decisions about the labor, inventory management, supply and production planning, freight and logistics and many other areas.

In Notes from the AI Frontier, McKinsey & Company highlight that, a 10 to 20% improvement in retail supply chain forecasting accuracy is likely to produce a 5% reduction in inventory costs and a 2 to 3% increase in revenues. Traditional supply chain forecasting tools have failed to deliver the desired results. With claims of industry-average inaccuracies of 32% in retailer supply chain demand forecasting, the potential impact of even modest forecasting improvements is immense for most retailers. As a result, many organizations are moving away from pre-packaged forecasting solutions, exploring ways to bring demand forecasting skills in-house and revisiting past practices which compromised forecast accuracy for computational efficiency.

A key focus of these efforts is the generation of forecasts at a finer level of temporal and (location/product) hierarchical granularity. Fine-grain demand forecasts have the potential to capture the patterns that influence demand closer to the level at which that demand must be met. Whereas in the past a retailer might have predicted short-term demand for a class of products at a market level or distribution level, for a month or week period, and then used the forecasted values to allocate units of a specific product in that class should be placed in a given store and day, fine-grain demand forecasting allows forecasters to build more localized models that reflect the dynamics of that specific product in a particular location.

Fine-grain Demand Forecasting Comes with Challenges

As exciting as fine-grain demand forecasting sounds, it comes with many challenges. First, by moving away from aggregate forecasts, the number of forecasting models and predictions which must be generated explodes. The level of processing required is either unattainable by existing forecasting tools, or it greatly exceeds the service windows for this information to be useful. This limitation leads to companies making tradeoffs in the number of categories being processed, or the level of grain in the analysis.

As examined in a prior blog post, Apache Spark can be employed to overcome this challenge, allowing modelers to parallelize the work for timely, efficient execution. When deployed on cloud-native platforms such as Databricks, computational resources can be quickly allocated and then released, keeping the cost of this work within budget.
The second and more difficult challenge to overcome is understanding that demand patterns that exist in aggregate may not be present when examining data at a finer level of granularity. To paraphrase Aristotle, the whole may often be greater than the sum of its parts. As we move to lower levels of detail in our analysis, patterns more easily modeled at higher levels of granularity may no longer be reliably present, making the generation of forecasts with techniques applicable at higher levels more challenging. This problem within the context of forecasting is noted by many practitioners going all the way back to Henri Theil in the 1950s.

As we move closer to the transaction level of granularity, we also need to consider the external causal factors that influence individual customer demand and purchase decisions. In aggregate, these may be reflected in the averages, trends and seasonality that make up a time series but at finer levels of granularity, we may need to incorporate these directly into our forecasting models.

Finally, moving to a finer level of granularity increases the likelihood the structure of our data will not allow for the use of traditional forecasting techniques. The closer we move to the transaction grain, the higher the likelihood we will need to address periods of inactivity in our data. At this level of granularity, our dependent variables, especially when dealing with count data such as units sold, may take on a skewed distribution that’s not amenable to simple transformations and which may require the use of forecasting techniques outside the comfort zone of many Data Scientists.

Accessing the Historical Data

See the Data Preparation notebook for details.

In order to examine these challenges, we will leverage public trip history data from the New York City Bike Share program, also known as Citi Bike NYC. Citi Bike NYC is a company that promises to help people, “Unlock a Bike. Unlock New York.” Their service allows people to go to any of over 850 various rental locations throughout the NYC area and rent bikes. The company has an inventory of over 13,000 bikes with plans to increase the number to 40,000. Citi Bike has well over 100,000 subscribers who make nearly 14,000 rides per day.

Citi Bike NYC reallocates bikes from where they were left to where they anticipate future demand. Citi Bike NYC has a challenge that is similar to what retailers and consumer goods companies deal with on a daily basis. How do we best predict demand to allocate resources to the right areas? If we underestimate demand, we miss revenue opportunities and potentially hurt customer sentiment. If we overestimate demand, we have excess bike inventory being unused.

This publicly available dataset provides information on each bicycle rental from the end of the prior month all the way back to the inception of the program in mid-2013. The trip history data identifies the exact time a bicycle is rented from a specific rental station and the time that bicycle is returned to another rental station. If we treat stations in the Citi Bike NYC program as store locations and consider the initiation of a rental as a transaction, we have something closely approximating a long and detailed transaction history with which we can produce forecasts.

As part of this exercise, we will need to identify external factors to incorporate into our modeling efforts. We will leverage both holiday events as well as historical (and predicted) weather data as external influencers. For the holiday dataset, we will simply identify standard holidays from 2013 to present using the holidays library in Python. For the weather data, we will employ hourly extracts from Visual Crossing, a popular weather data aggregator.

Citi Bike NYC and Visual Crossing data sets have terms and conditions that prohibit our directly sharing of their data. Those wishing to recreate our results should visit the data providers’ websites, review their Terms & Conditions, and download their datasets to their environments in an appropriate manner. We will provide the data preparation logic required to transform these raw data assets into the data objects used in our analysis.

Examining the Transactional Data

See the Exploratory Analysis notebook for details.

As of January 2020, the Citi Bike NYC bike share program consists of 864 active stations operating in the New York City metropolitan area, primarily in Manhattan. In 2019 alone, a little over 4-million unique rentals were initiated by customers with as many as nearly 14,000 rentals taking place on peak days.

Sample sales data visualization for Citi Bike NYC, illustrating demand across available rental stations.

Since the start of the program, we can see the number of rentals has increased year over year. Some of this growth is likely due to the increased utilization of the bicycles, but much of it seems to be aligned with the expansion of the overall station network.

Data visualization depicting increasing Citibike NYC rental demand and utilization between 2013 and 2020.
Data visualization depicting increasing Citibike NYC rental demand and utilization between 2013 and 2020.

Normalizing rentals by the number of active stations in the network shows that growth in ridership on a per-station basis has been slowly ticking up for the last few years in what we might consider to be a slight linear upward trend.

 Data visualization depicting increasing Citibike NYC per-station ridership between 2013 and 2020

Using this normalized value for rentals, ridership seems to follow a distinctly seasonal pattern, rising in the Spring, Summer and Fall and then dropping in Winter as the weather outside becomes less conducive to bike riding

Data visualization depicting the seasonality of Citibike NYC rental demand for years 2013 through 2020

his pattern appears to closely follow patterns in the maximum temperatures (in degrees Fahrenheit) for the city.

Data visualization showing the correlation between higher temperatures and increased demand Citibike NYC bike rentals.

While it can be hard to separate monthly ridership from patterns in temperatures, rainfall (in average monthly inches) does not mirror these patterns quite so readily

Data visualization for monthly average rainfall illustrates the difficulty in correlating weather to Citibike NYC rental demand.

Examining weekly patterns of ridership with Sunday identified as 1 and Saturday identified as 7, it would appear that New Yorkers are using the bicycles as commuter devices, a pattern seen in many other bike share programs.

Data visualization of Citibike NYC bike ridership by day of week indicates a pattern of commuter utilization seen with many other bike share programs.

Breaking down these ridership patterns by hour of the day, we see distinct weekday patterns where ridership spikes during standard commute hours. On the weekends, patterns indicate more leisurely utilization of the program, supporting our earlier hypothesis.

Data visualization of Citibike NYC bike ridership by hour of day displays rental activity occurring at all hours of the day and night.

An interesting pattern is that holidays, regardless of their day of week, show consumption patterns that roughly mimic weekend usage patterns. The infrequent occurrence of holidays may be the cause of erraticism of these trends. Still, the chart seems to support that the identification of holidays is important to producing a reliable forecast.

Data visualization of Citibike NYC bike holiday ridership by hour of day indicates a utilization pattern that roughly mimics weekend usage.

In aggregate, the hourly data appear to show that New York City is truly the city that never sleeps. In reality, there are many stations for which there are a large proportion of hours during which no bicycles are rented.

Data visualization of Citibike NYC bike ridership by number of individual stations that record 1 hour of inactivity during the day illustrates difficulty in using traditional methods to forecast demand.

These gaps in activity can be problematic when attempting to generate a forecast. By moving from 1-hour to 4-hour intervals, the number of periods within which individual stations experience no rental activity drops considerably though there are still many stations that are inactive across this timeframe.

Data visualization of Citibike NYC bike ridership by number of individual stations that record 4 hours of inactivity during the day illustrates the difficulty in using traditional methods to forecast demand.

Instead of ducking the problem of inactive periods by moving towards even higher-levels of granularity, we will attempt to make a forecast at the hourly level, exploring how an alternative forecasting technique may help us deal with this dataset. As forecasting for stations that are largely inactive isn’t terribly interesting, we’ll limit our analysis to the top 200 most active stations.

Forecasting Bike Share Rentals with Facebook Prophet

In an initial attempt to forecast bike rentals at the per-station level, we made use of Facebook Prophet, a popular Python library for time series forecasting. The model was configured to explore a linear growth pattern with daily, weekly and yearly seasonal patterns. Periods in the dataset associated with holidays were also identified so that anomalous behavior on these dates would not affect the average, trend and seasonal patterns detected by the algorithm.

Using the scale-out pattern documented in the previously referenced blog post, models were trained for most active 200 stations and 36-hour forecasts were generated for each. Collectively, the models had a Root Mean Squared Error (RMSE) of 5.44 with a Mean Average Proportional Error (MAPE) of 0.73. (Zero-value actuals were adjusted to 1 for the MAPE calculation.)

These metrics indicate that the models do a reasonably good job of predicting rentals but are missing when hourly rental rates move higher. Visualizing sales data for individual stations, you can see this graphically such as in this chart for Station 518, E 39 St & 2 Ave, which has an RMSE of 4.58 and a MAPE of 0.69:

See the Time Series notebook for details.

Data visualization showing the limitations of using a Facebook Prophet forecasting model configured to explore a linear growth pattern to predict localized demand. The model does a reasonably good job of predicting rentals for an individual Citibike NYC rental station but starts to miss when hourly rental rates move higher.

The model was then adjusted to incorporate temperature and precipitation as regressors. Collectively, the resulting forecasts had a RMSE of 5.35 and a MAPE of 0.72. While a very slight improvement, the models are still having difficulty picking up on the large swings in ridership found at the station level, as demonstrated again by Station 518 which had an RMSE of 4.51 and a MAPE of 0.68:

See the Time Series with Regressors notebook for details.
 Data visualization of a Facebook Prophet forecasting model configured to explore a linear growth pattern, with adjustments to incorporate to weather as regressors. While a very slight improvement, the models are still having difficulty picking up on the large swings in ridership found at the Citibike NYC station level.

This pattern of difficulty modeling the higher values in both the time series models is typical of working with data having a Poisson distribution. In such a distribution, we will have a large number of values around an average with a long-tail of values above it. On the other side of the average, a floor of zero leaves the data skewed. Today, Facebook Prophet expects data to have a normal (Gaussian) distribution but plans for the incorporation of Poisson regression have been discussed.

Alternative Approaches to Forecasting Supply Chain Demand

How might we then proceed with generating a forecast for these data? One solution, as the caretakers of Facebook Prophet are considering, is to leverage Poisson regression capabilities in the context of a traditional time series model. While this may be an excellent approach, it is not widely documented so tackling this on our own before considering other techniques may not be the best approach for our needs.

Another potential solution is to model the scale of non-zero values and the frequency of the occurrence of the zero-valued periods. The output of each model can then be combined to assemble a forecast. This method, known as Croston’s method, is supported by the recently released croston Python library while another data scientist has implemented his own function for it. Still, this is not a widely adopted method (despite the technique dating back to the 1970s) and our preference is to explore something a bit more out-of-the-box.

Given this preference, a random forest regressor would seem to make quite a bit of sense. Decision trees, in general, do not impose the same constraints on data distribution as many statistical methods. The range of values for the predicted variable is such that it may make sense to transform rentals using something like a square root transformation before training the model, but even then, we might see how well the algorithm performs without it.

To leverage this model, we’ll need to engineer a few features. It’s clear from the exploratory analysis that there are strong seasonal patterns in the data, both at the annual, weekly and daily levels. This leads us to extract year, month, day of week and hour of the day as features. We may also include a flag for holiday.

Using a random forest regressor and nothing but time-derived features, we arrive at an overall RMSE of 3.4 and MAPE of 0.39. For Station 518, the RMSE and MAPE values are 3.09 and 0.38, respectively:

See Temporal Notebook for details.

By leveraging precipitation and temperature data in combination with some of these same temporal features, we are able to better (though not perfectly) address some of the higher rental values. The RMSE for Station 518 drops to 2.14 and the MAPE to 0.26. Overall, the RMSE drops to 2.37 and MAPE to 0.26 indicating weather data is valuable in forecasting demand for bicycles.

See the Random Forest with Temporal & Weather Features notebook for details.

Data visualization of a Facebook Prophet forecasting model (for Citibike NYC) configured to explore a linear growth pattern, with adjustments using a random forest regressor and nothing but time-derived features.

Implications of the Results

Demand forecasting at finer levels of granularity may require us to think differently about our approach to modeling. External influencers which may be safely considered summarized in high-level time series patterns may need to be more explicitly incorporated into our models. Patterns in data distribution hidden at the aggregate level may become more readily exposed and necessitate changes in modeling approaches. In this dataset, these challenges were best addressed by the inclusion of hourly weather data and a shift away from traditional time series techniques towards an algorithm which makes fewer assumptions about our input data.

There may be many other external influencers and algorithms worth exploring, and as we go down this path, we may find that some of these work better for some subset of our data than for others. We may also find that as new data arrives, techniques that previously worked well may need to be abandoned and new techniques considered.

A common pattern we are seeing with customers exploring fine-grain demand forecasting is the evaluation of multiple techniques with each training and forecasting cycle, something we might describe as an automated model bake-off. In a bake-off round, the model producing the best results for a given subset of the data wins the round with each subset able to decide its own winning model type. In the end, we want to ensure we are performing good Data Science where our data is properly aligned with the algorithms we employ, but as is noted in article after article, there isn’t always just one solution to a problem and some may work better at one time than at others. The power of what we have available today with platforms like Apache Spark and Databricks is that we have access to the computational capacity to explore all these paths and deliver the best solution to our business.

--

Try Databricks for free. Get started today.

The post New Methods for Improving Supply Chain Demand Forecasting appeared first on Databricks.