Data Science, Machine Learning, Natural Language Processing, Text Analysis, Recommendation Engine, R, Python
Tuesday, 26 May 2020
Blockchain’s Disruptive Potential In An AV-reliant Post-Covid World
Read More on Datafloq
How AI Visual Inspection Systems Help Fight COVID-19
Read More on Datafloq
Need to tap Artificial Intelligence to fight Covid-19, says IT minister Ravi Shankar Prasad
RBI says no curbs in providing bank accounts to crypto traders
Monday, 25 May 2020
Read-only Jenkins Configuration
I’m excited to announce that the 'read-only' Jenkins feature is now available for preview. This feature allows restricting configuration UIs and APIs while providing access to essential Jenkins system configuration, diagnostics, and self-monitoring tools through Web UI. Such mode is critical for instances managed as code, e.g. with Jenkins Configuration-as-Code plugin. It is delivered as a part of the JEP-224: Readonly system configuration effort.
You will want to use at least Jenkins 2.238 to have all the features mentioned in this post.
Read-only Jenkins currently allows users to have access to:
-
job configuration
-
system configuration
-
plugin manager
-
system logs
-
cloud configuration
-
agent configuration
-
agent logs
For more planned integrations see the JENKINS-12548 epic.
Read-only Jenkins is split into three permissions:
-
Job/ExtendedRead - Read-only access to job configurations
-
existed since 2009 but the UI didn’t do anything to indicate to the users that they couldn’t edit the job configuration page. This has now been adapted to the new read-only engine.
-
-
Agent/ExtendedRead - Read-only access to agent configurations
-
existed since 2013 but it was undocumented and only allowed access to API and no UI
-
UI support added in Jenkins 2.238
-
-
Overall/SystemRead - System-wide read-only access. It is very useful for Jenkins instances managed as code, e.g. with help of the Jenkins Configuration as Code Plugin.
-
Introduced in Jenkins 2.222 as a part of JEP-224: Readonly system configuration
-
You can selectively grant the permission(s) as you wish.
Why do I want this?
Given the rise of the configuration-as-code plugin a lot of Jenkins instances are fully managed as code, which means that no changes are allowed through the UI.
The problem with this is you don’t know when new plugin versions are available and in order to see what other configuration options are available to a plugin you currently need the 'Administer' permission.
Read-only access to system administration information allows users who are not administrators to more easily debug build issues. For example, given a 'Jenkins' error message in a build the user can check:
-
which plugins are installed
-
the version of the plugin
This can allow the user to solve their issue themselves and makes it easier for the user to report an issue with a plugin directly to the maintainers.
What can I expect
All built in UI controls have been adapted to clearly distinguish between an editable control and a control you don’t have permission to edit:
Editable:

Non editable:

Note: there are other controls such as in the credentials and pipeline plugins that have not been updated yet.
Action buttons, (Such as 'Save' and 'Apply') have been hidden in most cases.
Work will continue on read-only configuration. Some plugins need support added and certain controls could have some improvements done to render better.
How can I use it?
These permissions are currently available in beta and for now disabled by default. You can enable them by installing the Extended read permission plugin v3.2 or above.
Then you will need to add the following permissions to a user / group depending on your use case:
-
Overall/SystemRead
-
Job/ExtendedRead
-
Agent/ExtendedRead
Note: You will need to set the Overall/Read and Job/Read permissions as well. You might want to consider creating a role containing the required permissions.
Here is an example using the Configuration as Code plugin and the Folder-based Authorization Strategy plugin:
jenkins:
authorizationStrategy:
folderBased:
globalRoles:
- name: "admin"
permissions:
- id: "Overall/Administer"
sids:
- "admin"
- name: "global read"
permissions:
- id: "Agent/ExtendedRead"
- id: "Overall/SystemRead"
- id: "Overall/Read"
- id: "Job/Read"
- id: "Job/ExtendedRead"
sids:
- "reader"I can’t see a configuration that I think should be allowed
Most of Jenkins itself has been updated to support read-only Jenkins, but not very many plugins. Please create an enhancement issue on the plugins issue tracker. If the plugin uses Jira to track issues, then you can add it to the JENKINS-12548 epic.
How do I update my plugin to support it
See the Read only view section of the developer documentation.
What’s next
In this release we introduce a foundation feature which is already supported in all key Jenkins core controls and in some plugins. There are many plugins which contribute to global configurations and diagnostics which still need to be adapted to support the new mode. We will keep working on this feature and its adoption so that the next LTS baseline in September provides a full-fledged user experience for Jenkins admins.
System read permission is a featured project in the UI/UX Hackfest happening May 25-29 2020. If you want to get involved please check it out!
Why is Customer Experience the New Marketing Strategy?
Strategy?Many businesses work day in day out to devise strategies that could help them to get better traction in their businesses. However, businesses pick different aspects when it comes to giving preference. Nowadays, many businesses are giving importance to devise a strategy where customers are given the first preference. Introducing customer experience into your marketing strategy has various implications. But before going to that, it is necessary to know what all reasons are making customer experience play such an important role while deciding business strategy.Despite the various decision-making metrics, It has been found in a Forrester study commissioned by Adobe that 80% of the businesses are trying to work on their customer experience.
Source: AdobeIt clearly shows that working on customer experience is going to help businesses improve results. And companies are obsessed with customers today, and that is one reason they are making mistakes as well in their marketing strategies.Companies are putting efforts that are going beyond the needed. Customer retention is directly related to the customer experience, and in certain circumstances, when retaining existing customers becomes more important than acquiring a ...
Read More on Datafloq
Blockchain Technology in Translation Processes
Read More on Datafloq
View: The need to look at India's technology agenda in a holistic way
Sunday, 24 May 2020
Will COVID-19 Surveillance Cause Citizens to Double Down on Online Privacy?
As concerning as this development is, other privacy concerns are emerging. They have not received as much attention, but they are potentially even more concerning. The COVID-19 pandemic has forced a growing number of people to move all of their communications online. As a result, their messages are subject to being intercepted by hackers.
Both of these concerns are causing citizens around the world to be more concerned than ever about their online privacy. They may decide to start taking new precautions to protect their digital information from hackers and unnecessary surveillance.
Polls show a growing concern about digital privacy
Pew Research recently released the findings of a new poll on digital privacy concerns during the coronavirus pandemic. Participants were informed that the United States and other governments were considering placing tracking apps on people’s mobile devices to monitor their locations. They were told that the geolocation data could presumably be ...
Read More on Datafloq
Saturday, 23 May 2020
India likely to use drones to beat back locusts
Jeff Bezos's family office bets on Sujal Patel's startup that is mapping human proteins
Friday, 22 May 2020
What Are The Pros & Cons of Modern ESBs
Read More on Datafloq
Adoption Picking Up for AI in Cybersecurity; More Skilled Humans Needed Too
By AI Trends Staff
AI is increasingly being put to use in the technology stacks of cybersecurity companies, but not at the expense of human experts who guide the rollout and work alongside the smart tools.
Before 2019, one in five cybersecurity software and service providers were employing AI, according to a study last year by Capgemini Research Institute, in a review of recent research published in DarkReading. Adoption was found to be “poised to skyrocket” by the end of 2020, with 63% of the firms planning to deploy AI in their solutions. Planned use in IT operations and the Internet of Things are predicted to see the most uptick.
Increased adoption of AI does not mean that security professionals on IT staffs are ready to hand off their responsibilities. A recent study conducted by White Hat Security at the RSA Conference 2020, held live at the end of February in San Francisco, found that 60% of security professionals are more confident when cyberthreat findings are verified by humans, over those generated by AI. One-third of respondents said intuition is the most important human element fueling analysis, while 21% said creativity is an advantage for humans.
Still, despite some reservations about AI, the White Hat survey found 70% of security professionals agreed that AI makes teams more efficient by taking over maybe 50% of the mundane tasks, freeing them for other work and reducing stress.
Some security professionals see their jobs as too complex to be taken over by machines, according to a recent Threat Intelligence report from the Ponemon Institute. Over half of the more than 1,000 IT professionals surveyed said they would not be able to train the AI to do the tasks their teams perform, and they are more qualified than AI to catch threats in real time. For protection of networks, close to half of respondents said human intervention was a necessity.
Nevertheless, the train has left the station for AI in cybersecurity. Some three-quarters of executives responding to the Cap Gemini survey said AI in cybersecurity speeds breach response, detection and remediation. Over 60% said AI also reduces the cost of detection and response.
Humans Said to Need the Help of AI in Cybersecurity
Humans need the help of AI to counter cybersecurity threats, suggests a recent report from KPMG and Oracle focused on trends in India. AI working with machine learning provides a powerful filter to sift through alerts and flag the most relevant, according to an account citing the report in The Hindu BusinessLine.
“Depending only on humans to counter the threat is no longer enough. It is far easier, efficient to keep track of different threat vectors and monitor an expanding threat surface with an AI-ML led approach,” stated Greg Jensen, Senior Principal Director of Security, Oracle. “Nearly all security providers now cite the use of some form of ML in their products as a means to protect against zero-day threats and malicious behaviors that evade more traditional forms of detection,” he added.
The Oracle KPMG Cloud Threat Report, based on a survey of 750 cybersecurity and IT professionals, found top priorities were the security of company financials and intellectual property. The respondents are using many products to combat threats, with 78% using more than 50 discrete cybersecurity products, and 37% using more than 100 products.
As IT organizations in India move more operations to the cloud, many are looking to define a cloud security strategy, which frequently employs a model of shared responsibility.
A shortage of skilled cyber security staff is a challenge for AI adoption in India, as it is globally, with not enough analysts available to triage alerts. AI is seen as being able to assist existing analysts in hunting and analyzing chains of attack.
Over 90% of the KPMG-Oracle survey respondents acknowledged the gap between the current cloud strategies and their ability to provide effective security and privacy controls. Oracle positions to help prescribe more intelligent automation of cybersecurity incorporating AI in response.
Unsupervised Machine Learning Seen as Effective
Machine learning models come in these different forms: Supervised, Reinforcement, Unsupervised and Semi-Supervised (also known as Active Learning). A recent account in Technative gives the nod to Unsupervised machine learning as the preference for cybersecurity.
Supervised Learning relies on a process of labeling in order to “understand” information. The machine learns from labeling lots of data and is able to “recognize” something only after someone, most likely a security professional, has already labeled it. The model cannot do it on its own, according to the author Ana Mezic of MixMode, a company offering a predictive threat modeling security service.
It is not usually the case in cybersecurity that you know exactly what you are looking for. If hackers use a method of attack that the security program has not seen before, the supervised machine learning system would not recognize it.
Unsupervised Learning draws inferences from datasets, searching for patterns out of the norm that could be dangerous. The software creates a baseline for a customer network, showing what a “normal day” looks like. A file transfer that is too large or sent at an odd time would be flagged. The model is optimized for predicting behavior, good enough that the company says it can detect zero-day attacks, those exploiting an unknown vulnerability.
Read the source articles in DarkReading, The Hindu BusinessLine and Technative.
Research Into Hardware Aims to Lower Demands and Expense of AI Software
By AI Trends Staff
With the energy and compute demands of AI machine learning models trending at what appears to be an unsustainable rate, researchers at Purdue University are experimenting with specialized hardware aimed at offloading some of the AI demands on software.
The approach exploits features of quantum computing, especially proton transport.
“Software is taking on most of the challenges in AI. If you could incorporate intelligence into the circuit components in addition to what is happening in software, you could do things that simply cannot be done today,” stated Shriram Ramanathan, a professor of materials engineering at Purdue University, in an account from Purdue University published on sciencesprings.
The reliance on software with massive energy needs to make AI work is not sustainable, suggested Ramanathan. If hardware and software could share intelligence features, the silicon might be able to achieve more with a given input of energy.
The hardware the Purdue team is developing is made of quantum material, with properties the team is working to understand and apply to solving problems in electronics. The way software uses tree-like memory to organize information into various “branches,” is how the human brain categorizes information and makes decisions.
“Humans memorize things in a tree structure of categories. We memorize ‘apple’ under the category of ‘fruit’ and ‘elephant’ under the category of ‘animal,’ for example,” said Hai-Tian Zhang, a Lillian Gilbreth postdoctoral fellow in Purdue’s College of Engineering. “Mimicking these features in hardware is potentially interesting for brain-inspired computing.”
The team introduced a proton to a quantum material called neodymium nickel oxide. They discovered that applying an electric pulse to the material moves around the proton. Each new position of the proton creates a memory state; multiple electric pulses create a branch made up of memory states.
“We can build up many thousands of memory states in the material by taking advantage of quantum mechanical effects. The material stays the same. We are simply shuffling around protons,” Ramanathan said.
The team showed that the material is capable of learning the numbers 0 through 9, a baseline test of AI. The demonstration of these tasks at room temperature in a material is a step toward showing that hardware could offload tasks from software.
“This discovery opens up new frontiers for AI that have been largely ignored because implementing this kind of intelligence into electronic hardware didn’t exist,” Ramanathan said.
The results of this study are published in the journal Nature Communications.
Whether AI machine learning is making unreasonable power demands is also being investigated by John Naughton, professor of the public understanding of technology at the Open University, a public research university, the largest university in the UK for undergraduate education. He is the author of “From Gutenberg to Zuckerberg: What You Really Need to Know About the Internet.”
Researchers at Nvidia, the manufacturer of GPUs (graphic processing units) now used in most machine-learning systems, developed a natural language model that was 24 times bigger than its predecessor, and 34 percent better at its learning task, Naughton wrote in a recent account in The Guardian: Training the final model took 512 V100 GPUs running continuously for 9.2 days. One expert calculated that would be three times the yearly energy consumption of the average American.
“You don’t have to be Einstein to realize that machine learning can’t continue on its present path, especially given the industry’s frenetic assurances that tech giants are heading for an ‘AI everywhere’ future,” Naughton stated.
He, too, suggests that advances in hardware could help make the demands of AI more practical, noting that the new Apple iPhone 11 includes Apple’s A13 chip, which incorporates neural network software behind recent advances in natural language and image processing.
IBM Experimenting with The Neural Computer
Meanwhile at IBM, research is going on into the “neural computer,” a new type of computer designed to develop AI algorithms and assist in computational neuroscience. The Neural Computer is a deep “neuroevolution” system that combines the hardware implementation of an Atari 2600, image preprocessing, and AI algorithms in an optimized pipeline, according to a recent account in VentureBeat. (The Atari 2600, originally branded as the Atari Video Computer Systems, was introduced in 1977.)
The research team reports in an IBM technical research paper released earlier this year, that results so far have achieved a record training time of 1.2 million image frames per second.
Video games are a well-established platform for AI and machine learning research. In certain domains like reinforcement learning, the AI learns optimal behaviors by interacting with the environment in pursuit of rewards, such as good game scores. AI algorithms developed within games have been shown to be adaptable to practical uses, including protein folding prediction. If the results from IBM’s Neural Computer prove to be repeatable, the system could be used to accelerate the development of AI algorithms.
Over the course of five experiments, IBM researchers ran 59 Atari 2600 games on the Neural Computer. It required 6 billion game frames in total and failed at challenging exploration games like Montezuma’s Revenge and Pitfall. But it managed to outperform a popular baseline — a Deep Q-network, an architecture pioneered by DeepMind — in 30 out of 59 games after 6 minutes of training (200 million training frames). This compares to the Deep-Q network’s 10 days of training. With 6 billion training frames, it surpassed the Deep Q-network in 36 games while taking 2 orders of magnitude less training time (2 hours and 30 minutes).
Read the source articles in sciencesprings, The Guardian, VentureBeat and at IBM Research.
Marty the Robot Rolls out AI in the Supermarket
By John P. Desmond, AI Trends Editor
When six-foot-four inch Marty first rolled into Stop & Shop, the robot walked into history. Social robot experts say it is among the first instance of a robot deployed in a customer environment, namely supermarkets in the Northeast.
Marty rolls around the store looking for spills with its three cameras. It does take the place of the human worker, called an associate, that did the same thing, but it means the associate can do something else. Doing the walk-around of the store is seen as a mundane task.
Marty does not talk or tell jokes. Unlike Alexa, who many children in the store undoubtedly interact with at home, Marty will not respond. The robot does notify associates when it sees with its computer vision that something on the floor needs to be cleaned up, through the public address system. An associate comes over to clean it up, and presses a button on Marty that it’s done. Marty takes a picture of the cleaned-up aisle.
Badger Technologies Rolled Out 500 Martys in December
The AI in Marty is concentrated on the machine vision and the collision-avoidance navigation features, according to Richard Rowland, CEO of Badger Technologies, makers of Marty. Last December, after a year of trials, Badger rolled out 500 multi-purpose robots into Stop & Shop and Giant/Martin’s grocery stores on the East Coast. Each Marty is equipped with navigation systems, high-resolution cameras, many sensors and its software systems.
When Marty gets to the store, it takes a trip in a shopping cart for 45 minutes to an hour to map the floor. “It produces a 3D map that it stores in its memory as its reference point,” Rowland said in an interview with AI Trends. The robot is instructed to “pose” at the top and bottom of each aisle as it makes its run around the store. At each pose, the three cameras on the right side of the robot are turned to look down the aisle and take pictures. It can see up for 70 feet, with its three cameras focused on close, medium and far distances.
A primary design goal for Marty was to travel safely in the stores. “We were one of the first to introduce a robot into such a public setting,” Rowland said. “We would rather be working in a warehouse than out in the middle of many shoppers, so the primary goal was to operate safely with a lot of foot traffic around and to not bump into anything.”
Beyond that, the robot was designed to scan the floor to see if any spills need to be cleaned up. “The imaging function is the second major area of AI on the robot,” Rowland said.
The designers wanted the robot to be “in the background” as much as possible, so they colored it a muted gray. It gives off an audible beep to let shoppers know it’s in the area, and it has traveling lights. It weighs 140 pounds and costs $35,000.
“The application of the ‘googly eye’ was somewhat of an accident,” Rowland said. One of the Stop & Shop associates took it upon herself to try to make the robot look more friendly, so she put some eyes on it. Now Marty has a social media following. On Saturday, Jan. 12, 325 Stop & Shop locations celebrated a one-year birthday for Marty, inviting kids in for refreshments and coloring books. Stories about Marty parties ran in many local newspapers and on TV news. Many videos of Marty are available on YouTube.
Marty Anthropomorphized
Anthropomorphizing is assigning of human traits, emotions or intentions to non-human entities. It is considered to be an innate tendency of human psychology. So by virtue of the Googly Eyes, voila! Marty is anthropomorphized. This has been a learning experience for Badger, which has not so far employed any industrial anthropologists to guide the introduction of Marty.
“I’d love to tell you it was intentional and that we did studies, but it was a reaction on our part. It happened accidentally. And we have learned from that. It takes the edge off an upright piece of technology moving around a supermarket,” Rowland said.
The public embracing of Marty is making a difference. “In hindsight, it did achieve the acceptance of a piece of high tech equipment more than it might have been. It softened the change of dynamics,” Rowland said.
The robot’s height and camera placement enables it to scan upper shelves to spot out-of-stock or mis-priced products. Six stores were testing this capability when we spoke in early May.
Marty A First
Dr. John Sullins of Sonoma State University, Calif., an expert on computer ethics and the philosophical implications of technologies including Robotics and AI, has a range of experience with robotics in the workplace. “Marty is one of the first robotic systems we have seen that is operating autonomously in the same environment with customers,” Prof. Sullins said in an interview with AI Trends. “On an assembly line, robots are commonly fenced off or otherwise physically separated from other workers for safety purposes. Systems like Marty have to be designed to move around a crowded aisle plotting a path to avoid customers and not rudely entering their physical space. As humans, we do that sort of thing unconsciously, but for a machine to do something similar takes a lot of programming skill and years of product trials before it can be done safely.”
The release of Marty into the “wild” of a supermarket is a green field for social robot researchers. Dr. Sullins, for example, is very involved with developing robot ethical standards within IEEE. One area of study is the setting of expectations between humans and the robots. “Robot ethics requires transparency in human robot interactions,” Dr. Sullins said. Marty’s job in the supermarket is not necessarily spelled out to shoppers, leading to speculation about what the robot is doing. “We need to convey that kind of information quickly to the general public as it interacts with machines like Marty,” Sullins suggested.
Some of Marty’s behavior Dr. Sullins would classify as “nudging,” actions of intelligent, robotic systems designed to influence behavior, such as moving pedestrian traffic away from spills and directing human workers to attend to the spills. “More subtle nudges are the googly eyes, the Easter Bunny decorations and Covid face mask, which are designed to make people feel more comfortable around the machine,” he said.
[Ed. Note: Dr. Sullins and his compatriots in the IEEE are looking for AI professionals interested in crafting standards for the ethical use of intelligent autonomous systems. Dr. Sullins is co-chair of P7008 – Standard for Ethically Driven Nudging for Robotic, Intelligent and Autonomous Systems (https://standards.ieee.org/project/7008.html). Dr. Sullins sees Marty demonstrating some nudging behavior.]
Badger: an Entrepreneurial Success Story
The early version of Marty had a tablet computer mounted on the side that the robot designers thought could help customers locate products, and essentially open the door to customers communicating directly with the supermarket robot. But the timing was not right and that plan was shelved. “The early conclusion was that it distracted it from the main job,” Rowland said. “A future application could be verbal; that makes sense. It could be a future requirement.”
Marty has been good for Badger Technologies, which Rowland founded in 2017. He had spent 28 years in various engineering capacities at Lexmark, the printer company located in Lexington, Kentucky. He led an innovation project team focused on retail robotics; the team selected Jabil to do the advanced manufacturing required. The first live demo was conducted in 2016. Soon after, Lexmark management decided to continue its focus on printers and imaging products, so Rowland and a 12-person team formed Badger.
Jabil helped Badger in a joint development effort, enabling the team to move into retail in the spring of 2017. By late summer of that year, Jabil acquired Badger Technologies as an independent product division of the global manufacturing company’s operations. Jabil describes itself as a “manufacturing solutions provider that delivers comprehensive design, manufacturing, supply chain and product management services.” Jabil was founded in 1966 and is based in St. Petersburg, Fla.
In early 2018, Badger began a pilot robotics rollout with Ahold Delhaize USA Brands, owners of Stop & Shop and Giant supermarkets. The supermarket business is tough, with margins hovering in the one to two percent range. Ahold Delhaize of Zaandam, Netherlands, was formed with a merger of two companies in 2016, with the idea of squeezing out some savings and diversifying from conventional retail brands. The Ahold Delhaize USA division oversees Food Lion, Giant Food, Giant/Martin’s, Hannaford, and Stop & Shop, as well as e-grocer Peapod; Retail Business Services, a U.S. support services company providing services to the brands; and Peapod Digital Labs, its e-commerce engine. Operating more than 2,000 stores across 23 states, Ahold Delhaize USA is No. 4 on Progressive Grocer’s 2019 Super 50 list of the top grocers in the United States.
Labor issues hit Stop & Shop in 2019, with 30,000 workers going on strike for 11 days, costing the company between $90 million and $100 million, according to an account in The New Food Economy The strike was resolved, the workers came back, and Marty still patrols the aisles, having become well-established and friendlier.
Exploring AI Dependence Upon ‘Artificial Stupidity’ For Autonomous Cars
By Lance Eliot, the AI Trends Insider
We all generally seem to know what it means to say that someone is intelligent.
In contrast, when you label someone as “stupid,” the question arises as to what exactly that means. For example, does stupidity imply the lack of intelligence in a zero-sum fashion, or does stupidity occupy its own space and sit adjacent to intelligence as a parallel equal?
Let’s do a thought experiment on this weighty matter.
Suppose we somehow had a bucket filled with intelligence. We are going to pretend that intelligence is akin to something tangible and that we can essentially pour it into and possibly out of a bucket that we happen to have handy. Upon pouring this bucket filled with intelligence onto say the floor, what do you have left?
One answer is that the bucket is now entirely empty and there is nothing left inside the bucket at all. The bucket has become vacuous and contains absolutely nothing. Another answer is that the bucket upon being emptied of intelligence has a leftover that consists of stupidity. In other words, once you’ve removed so-called intelligence, the thing that you have remaining is stupidity.
I realize this is a seemingly esoteric discussion but, in a moment, you’ll see that the point being made has a rather significant ramification for many important things, including and particularly for the development and rise of Artificial Intelligence (AI).
Can intelligence exist without stupidity, or in a practical sense is there always some amount of stupidity that must exist if there is also the existence of stupidity?
Some assert that intelligence and stupidity are Zen-like yin and yang. In this perspective, you cannot grasp the nature of intelligence unless you also have a semblance of stupidity as a kind of measuring stick.
It is said that humans become increasingly intelligent over time, and thus are reducing their levels of stupidity. You might suggest that intelligence and stupidity are playing a zero-sum game, namely that as your intelligence rises you are simultaneously reducing your level of stupidity (similarly, if your stupidity rises, this implies that your intelligence lowers).
Can humans arrive at a 100% intelligence and a zero amount of stupidity, or are we fated to always have some amount of stupidity, no matter how hard we might try to become fully intelligent?
Returning to the bucket metaphor, some would claim that there will never be the case that you are completely and exclusively intelligent and have expunged stupidity. There will always be some amount of stupidity that’s sitting in that bucket.
If you are clever and try hard, you might be able to narrow down how much stupidity you have, though there is still some amount of stupidity in that bucket.
Does having stupidity help intelligence or is it harmful to intelligence?
You might be tempted to assume that any amount of stupidity is a bad thing and therefore we must always be striving to keep it caged or otherwise avoid its appearance. But we need to ask whether that simplistic view of tossing stupidity into the “bad” category and placing intelligence into the “good” category is potentially missing something more complex. You could argue that by being stupid, at times, in limited ways, doing so offers a means for intelligence to get even better.
When you were a child, suppose you stupidly tripped over your own feet, and after doing so, you came to the realization that you were not carefully lifting your feet. Henceforth, you became more mindful of how to walk and thus became intelligent at the act of walking. Maybe later in life, while walking on a thin curb, you managed to save yourself from falling off the edge of the curb, partially due to the earlier in life lesson that was sparked by stupidity and became part of your intelligence.
Of course, stupidity can also get us into trouble.
Despite having learned via stupidity to be careful as you walk, one day you decide to strut on the edge of the Grand Canyon. While doing so, oops, you fall off and plunge into the chasm.
Was it an intelligent act to perch yourself on the edge like that? Apparently not.
As such, we might want to note that stupidity can be a friend or a foe, and it is up to the intelligence portion to figure out which is which in any given circumstance and any given moment.
You might envision that there is an eternal struggle going on between the intelligence side and the stupidity side.
On the other hand, you might equally envision that the intelligence side and stupidity side are pals, each of which tugs at the other, and therefore it is not especially a fight as it is a delicate dance and form of tension about which should prevail (at times) and how they can each moderate or even aid the other.
This preamble provides a foundation to discuss something increasingly becoming worthy of attention, namely the role of Artificial Intelligence and (surprisingly) the role of Artificial Stupidity.
For my indication of the grand convergence that has led to today’s AI, see this link: https://aitrends.com/ai-insider/grand-convergence-explains-rise-self-driving-cars/
For the importance of AI having self-awareness, see my article here: https://aitrends.com/ai-insider/self-awareness-self-driving-cars-know-thyself/
For why it is crucial to have AI algorithmic transparency, see my review here: https://aitrends.com/ai-insider/algorithmic-transparency-self-driving-cars-call-action/
For my assessing whether AI can have the motivation, see the article here: https://aitrends.com/ai-insider/motivational-ai-bounded-irrationality-self-driving-cars/
Thinking Seriously About Artificial Stupidity
We hear every day about how our lives are being changed via the advent of Artificial Intelligence.
AI is being infused into our smartphones, and into our refrigerators, and into our cars, and so on.
If we are intending to place AI into the things we use, it begs the question as to whether we need to consider the yang of the yin, specifically do we need to be cognizant of Artificial Stupidity?
Most people snicker upon hearing or seeing the phrase “Artificial Stupidly,” and they assume it must be some kind of insider joke to refer to such a thing.
Admittedly, the conjoining of the words artificial and stupidity seems, well, perhaps stupid in of itself.
But, by going back to the earlier discussion about the role of intelligence and the role of stupidity as it exists in humans, you can recast your viewpoint and likely see that whenever you carry on a discussion about intelligence, one way or another you inevitably need to also be considering the role of stupidity.
Some suggest that we ought to use another way of expressing Artificial Stupidity to lessen the amount of snickering that happens. Floated phrases include Artificial Unintelligence, Artificial Humanity, Artificial Dumbness, and others, none of which have caught hold as yet.
Please bear with me and accept the phrasing of Artificial Stupidity and also go along with the belief that it isn’t stupid to be discussing Artificial Stupidity.
Indeed, you could make the case that the act of not discussing Artificial Stupidity is the stupid approach since you are unwilling or unaccepting of the realization that stupidity exists in the real world and therefore in the artificial world of computer systems that are we attempting to recreate intelligence, you would be ignoring or blind to what is essentially the other half of the overall equation.
In short, some say that true Artificial Intelligence requires a combination of the “smart” or good AI that we think of today and the inclusion of Artificial Stupidity (warts and all), though the inclusion must be done in a smart way.
Indeed, let’s deal with the immediate knee jerk reaction that many have of this notion by dispelling the argument that by including Artificial Stupidity into Artificial Intelligence you are inherently and irrevocably introducing stupidity and presumably, therefore, aiming to make AI stupid.
Sure, if you stupidly add stupidity, you have a solid chance of undermining the AI and rendering it stupid.
On the other hand, in recognition of how humans operate, the inclusion of stupidity, when done thoughtfully, could ultimately aid the AI (think about the story of tripping over your own feet as a child).
Here’s something that might really get your goat.
Perhaps the only means to achieve true and full AI, which is not anywhere near to human intelligence levels to-date, consists of infusing Artificial Stupidity into AI; thus, as long as we keep Artificial Stupidity at arm’s length or as a pariah, we trap ourselves into never reaching the nirvana of utter and complete AI that is able to seemingly be as intelligent as humans are.
Ouch, by excluding Artificial Stupidity from our thinking, we might be damming ourselves to not arriving at the pinnacle of AI.
That’s a punch to the gut and so counterintuitive that it often stops people in their tracks.
There are emerging signs that the significance of revealing and harnessing artificial stupidity (or whatever it ought to be called), can be quite useful.
One such area, I assert, involves the inclusion of artificial stupidity into the advent of true self-driving driverless autonomous cars.
Shocking?
Maybe so.
Let’s unpack the matter.
For my framework about AI self-driving autonomous cars, see this link: https://aitrends.com/ai-insider/framework-ai-self-driving-driverless-cars-big-picture/
On the dangers of AI becoming a Frankenstein, see my analysis: https://aitrends.com/ai-insider/frankenstein-and-ai-self-driving-cars/
To understand the cognitive elements of autonomous cars, see my explanation here: https://aitrends.com/ai-insider/cognitive-timing-for-ai-self-driving-cars/
Exploiting Artificial Stupidity For Gain
When referring to true self-driving cars, I’m focusing on Level 4 and Level 5 of the standard scale used to gauge autonomous cars. These are self-driving cars that have an AI system doing the driving and there is no need and typically no provision for a human driver.
The AI does all the driving and any and all occupants are considered passengers.
On the topic of Artificial Stupidity, it is worthwhile to quickly review the history of how the terminology came about.
In the 1950s, the famous mathematician and pioneering computer scientist Alan Turing proposed what has become known as the Turing test for AI.
Simply stated, if you were presented with a situation whereby you could interact with a computer system imbued with AI, and at the same time separately interact with a human too, and you weren’t told beforehand which was which (let’s assume they are both hidden from view), upon your making inquiries of each, you are tasked with deciding which one is the AI and which one is the human.
We could then declare the AI a winner as exhibiting intelligence if you could not distinguish between the two contestants. In that sense, the AI is indistinguishable from the human contestant and must ergo be considered equal in intelligent interaction.
There is a twist to the original Turing test that many don’t know about.
One qualm expressed was that you might be smarmy and ask the two contestants to calculate say pi to the thousandth digit.
Presumably, the AI would do so wonderfully and readily tell you the answer in the blink of an eye, doing so precisely and abundantly correctly. Meanwhile, the human would struggle to do so, taking quite a while to answer if using paper and pencil to make the laborious calculation, and ultimately would be likely to introduce errors into the answer.
Turing realized this aspect and acknowledged that the AI could be essentially unmasked by asking such arithmetic questions.
He then took the added step, one that some believe opened a Pandora’s box, and suggested that the AI ought to avoid giving the right answers to arithmetic problems.
In short, the AI could try to fool the inquirer by appearing to answer as a human might, including incorporating errors into the answers given and perhaps taking the same length of time that doing the calculations by hand would take.
Starting in the early 1990s, a competition was launched that is akin to the Turing test, offering a modest cash prize and has become known as the Loebner Prize, and in this competition, the AI systems are typically infused with human-like errors to aid in fooling the inquirers into believing the AI is the human. There is controversy underlying this, but I won’t go into that herein. A now-classic article appeared in 1991 in The Economist about the competition.
Notice that once again we have a bit of irony that the introduction of stupidity is being done to essentially portray that something is intelligent.
This brief history lesson provides a handy launching pad for the next elements of this discussion.
Let’s boil down the topic of Artificial Stupidity into two main facets or definitions:
1) Artificial Stupidity is the purposeful incorporation of human-like stupidity into an AI system, doing so to make the AI seem more human-like, and being done not to improve the AI per se but instead to shape the perception of humans about the AI as being seemingly intelligent.
2) Artificial Stupidity is an acknowledgment of the myriad of human foibles and the potential inclusion of such “stupidity” into or alongside the AI in a conjoined manner that can potentially improve the AI when properly managed.
One common misnomer that I’d like to dispel about the first part of the definition involves a somewhat false assumption that the computer potentially is going to purposefully miscalculate something.
There are some that shriek in horror and disdain that there might be a suggestion that the computer would intentionally seek to incorrectly do a calculation, such as figuring out pi but doing so in a manner that is inaccurate.
That’s not what the definition necessarily implies.
It could be that the computer might correctly calculate pi to the thousandth digit, and then opt to tweak some of the digits, which it would say keep track of, and do this in a blink of the eye, and then wait to display the result after an equivalent of the human-by-hand amount of time.
In that manner, the computer has the correct answer internally and has only displayed something that seems to have errors.
Now, that certainly could be bad for the humans that are relying upon what the computer has reported but note that this is decidedly not the same as though the computer has in fact miscalculated the number.+
There’s more than can be said about such nuances, but for now, let’s continue forward.
Both of those variants of Artificial Stupidity can be applied to true self-driving cars.
Doing so carries a certain amount of angst and will be worthwhile to consider.
For my detailed review of the Turing Test, see this link: https://aitrends.com/ai-insider/turing-test-ai-self-driving-cars/
On the problems of probabilistic reasoning in AI, take a look at my indication: https://aitrends.com/ai-insider/probabilistic-reasoning-ai-self-driving-cars/
Common sense reasoning is an open-ended challenge and needs to be considered, see my article: https://aitrends.com/ai-insider/common-sense-reasoning-and-ai-self-driving-cars/
A controversial perspective is that perhaps we need to restart our understanding and approach to AI, see this discussed here: https://aitrends.com/ai-insider/starting-over-on-ai-and-self-driving-cars/
Artificial Stupidity And True Self-Driving Cars
Today’s self-driving cars that are being tried out on our public roadways have already gotten a reputation for their driving prowess. Overall, driverless cars to-date are akin to a novice teenage driver that is timid and somewhat hesitant about the driving task.
When you encounter a self-driving car, it will often try to create a large buffer zone between it and the car ahead, attempting to abide by the car lengths rule-of-thumb that you were taught when first learning to drive.
Human drivers generally don’t care about the car lengths safety zone and edge up on other cars, doing so to their own endangerment.
Here’s another example of driving practices.
Upon reaching a stop sign, a driverless car will usually come to a full and complete stop. It will wait to see that the coast is clear, and then cautiously proceed. I don’t know about you, but I can say that where I drive, nobody makes complete stops anymore at stop signs. A rolling stop is a norm nowadays.
You could assert that humans are driving in a reckless and somewhat stupid manner. By not having enough car lengths between your car and the car ahead, you are increasing your chances of a rear-end crash. By not fully stopping at a stop sign, you are increasing your risks of colliding with another car or a pedestrian.
In a Turing test manner, you could stand on the sidewalk and watch cars going past you, and by their driving behavior alone you could likely ascertain which are the self-driving cars and which are the human-driven cars.
Does that sound familiar?
It should, since this is roughly the same as the arithmetic precision issue earlier raised.
How to solve this?
One approach would be to introduce Artificial Stupidity as defined above.
First, you could have the on-board AI purposely shorten the car’s length buffer to appear as though it is driving in the same manner as humans. Likewise, the AI could be modified to roll through stop signs. This is all rather easily arranged.
Humans watching a driverless car and a human-driven car would no longer be able to discern one such car from the other since they both would be driving in the same error-laden way.
That seems to solve one problem as it relates to the perception that we humans might have about whether the AI of self-driving cars is intelligent or not.
But, wait for a second, aren’t we then making the AI into a riskier driver?
Do we want to replicate and promulgate this car-crash causing risky human driving behaviors?
Sensibly, no.
Thus, we ought to move to the second definitional portion of Artificial Stupidity, namely by incorporating these “stupid” ways of driving into the AI system in a substantive way that allows the AI to leverage those aspects when applicable and yet also be aware enough to avoid them or mitigate them when needed.
Rather than having the AI drive in human error-laden ways and do so blindly, the AI should be developed so that it is well-equipped enough to cope with human driving foibles, detecting those foibles and being a proper defensive driver, along with leveraging those foibles when the circumstances make sense to do so (for more on this, see my posting here).
On the pranking of AI autonomous cars, see my assessment here: https://aitrends.com/ai-insider/pranking-of-ai-self-driving-cars/
One outside-the-box approach to AI includes child-learning, see my recap at this link: https://www.aitrends.com/ai-insider/ai-machine-child-deep-learning-the-case-of-ai-self-driving-cars/
Applying these topics to one-shot learning is an intriguing opportunity, see my analysis: https://www.aitrends.com/ai-insider/seeking-one-shot-machine-learning-the-case-of-ai-self-driving-cars/
For my comments about the infamous AI paperclip problem, see the link here: https://aitrends.com/ai-insider/super-intelligent-ai-paperclip-maximizer-conundrum-and-ai-self-driving-cars/
Conclusion
One of the most unspoken secrets about today’s AI is that it does not have any semblance of common-sense reasoning and in no manner whatsoever has the capabilities of overall human reasoning (many refer to such AI as Artificial General Intelligence or AGI).
As such, some would suggest that today’s AI is closer to the Artificial Stupidity side of things than it is to the true Artificial Intelligence side of things.
If there is a duality of intelligence and stupidity in humans, presumably you will need a similar duality in an AI system if it is to be able to exhibit human intelligence (though, some say that AI might not have to be so duplicative).
On our roads today, we are unleashing so-called AI self-driving cars, yet the AI is not sentient and not anywhere close to being sentient.
Will self-driving cars only be successful if they can climb further up the intelligence ladder?
No one yet knows, and it’s certainly not a stupid question to be asked.
Copyright 2020 Dr. Lance Eliot
This content is originally posted on AI Trends.
[Ed. Note: For reader’s interested in Dr. Eliot’s ongoing business analyses about the advent of self-driving cars, see his online Forbes column: https://forbes.com/sites/lanceeliot/]
AI Is Watching You Work, With Mixed Results
By AI Trends Staff
Advances in AI and sensors are providing new ways to digitize manual labor, giving managers new insights and potentially new leverage on employees
Many jobs in manufacturing require a dexterity and creativity that robots and software are unlikely to match any time soon. However, manufacturing jobs are likely to change based on how it seems best to work with the AI going forward.
An experiment has been going on at an auto parts manufacturing plant in Battle Creek, Mich. since 2017 to capture worker movements all day long in the hopes of identifying bottlenecks in production. The plant of Denso, the global auto parts manufacturer, has been piping video into machine learning software from startup Drishti, according to a recent account in Wired.
“In the past, we would take a line that was struggling and bring a bunch of people down with stopwatches to try to make it better,” stated Tony Huffman says, a production supervisor at the plant. The Drishti system logs the “cycle time” for every worker all day, for every shift. Plant managers analyze the data to look for sometimes subtle bottlenecks. “Everything flows better and is smoother,” Huffman stated.
Denso originated as part of Toyota, which still owns a stake in the company, and like its parent company uses the kaizen philosophy of manufacturing, which encourages workers at all levels to participate in improving how a plant operates. In US plants, the prevailing culture may not be so conducive to allowing workers to have as much impact on the production process.
“In Japan and most other countries, there would be some collective way the workers could say ‘This is too fast.’ In the US plants there often isn’t,” stated Susan Helper, an economics professor at Case Western Reserve University, who studies manufacturing.
Denso’s US plants are not unionized, but Raja Shembekar, a vice president at Denso, stated that the company has a good relationship with all its workers, who he says buy into the program when they see it could be helpful to them.
Call Center Worker Skeptical of AI Software
The monitoring software needs to be very smart to pick up nuances in, for instance, a speaking voice. An example of this was published recently in The Verge, describing the experience of a call center worker they named “Angela” to protect her identity. Her workplace was monitored by software from Voci Technologies of Pittsburgh, a company offering call center performance metrics
Her other measurements were excellent, but the program consistently marked Angela down for expressing negative emotions. She found that perplexing; her human managers had previously praised her empathetic manner on the phone; no one could tell her exactly why she was getting penalized. Her best guess was that the AI was interpreting her fast-paced and loud speaking style, periods of silence (a result of trying to meet a metric meant to minimize putting people on hold), and expressions of concern as negative.
“It makes me wonder if it’s privileging fake empathy, sounding really chipper and being like, ‘Oh, I’m sorry you’re dealing with that,’” stated Angela, who asked to use a pseudonym out of fear of retribution. “Feeling like the only appropriate way to display emotion is the way that the computer says, it feels very limiting. It also seems to not be the best experience for the customer, because if they wanted to talk to a computer, then they would have stayed with IVR [Interactive Voice Response].”
A spokesperson for Voci said the company trained its machine learning model on thousands of hours of audio that crowdsourced workers labeled as demonstrating positive or negative emotions. He acknowledged that the evaluations are subjective but all together should control for variables, and that each call center decides how to use the data provided.
Angela is nervous about the company’s future automation plans for AI in the call center. One of them is from Cogito Corp. of Boston, which supplies “real time conversational guidance.” The product’s AI coaches workers in real time, suggesting speaking more slowly, or with more energy or to express empathy. One training video on the company website is titled, “The ROI of Empathy.”
Whether automated empathy has a future remains to be seen. Angela is concerned about being penned in. “If you automate everything, you lose the flexibility to have a human connection,” she stated.
Workers in Amazon warehouses are subject to all kinds of electronic monitoring, described in The Verge account, which generally rewards faster production and seems to have no regard for the need of human beings to take a break and relax sometimes.
“You’re not stopping,” said “Jake,” an Amazon warehouse worker also using a pseudonym. “You are literally not stopping. It’s like leaving your house and just running and not stopping for anything for 10 straight hours, just running.”
Maybe the software can be reprogrammed to detect an inhuman pace, and to keep workers happy while they get the job done. That would be quite a breakthrough.
Mynvax, IISc-incubated startup looks at Covid-19 vaccine in 18 months
Chatbots in the Industry- Setting Up a New Trend
Read More on Datafloq
Thursday, 21 May 2020
10 Secrets to Remote Work UX Designers Might be Missing
Read More on Datafloq
Apple-Google contact tracing tech draws interest in 23 countries, some hedge bets
Bharat Biotech signs pact with US university for Covid-19 vaccine
Chinese Blockchain Developments Will Fundamentally Alter the World Economy
Read More on Datafloq
Telemedicine collective StepOne makes it to Aarogya Setu Mitr
Why NLP Is a Promising Technology for Business
Read More on Datafloq
Wednesday, 20 May 2020
New Pandas UDFs and Python Type Hints in the Upcoming Release of Apache Spark 3.0™
Pandas user-defined functions (UDFs) are one of the most significant enhancements in Apache Spark for data science. They bring many benefits, such as enabling users to use Pandas APIs and improving performance.
However, Pandas UDFs have evolved organically over time, which has led to some inconsistencies and is creating confusion among users. The full release of Apache Spark 3.0, expected soon, will introduce a new interface for Pandas UDFs that leverages Python type hints to address the proliferation of Pandas UDF types and help them become more Pythonic and self-descriptive.
This blog post introduces new Pandas UDFs with Python type hints, and the new Pandas Function APIs including grouped map, map, and co-grouped map.
Pandas UDFs
Pandas UDFs were introduced in Spark 2.3, see also Introducing Pandas UDF for PySpark. Pandas is well known to data scientists and has seamless integrations with many Python libraries and packages such as NumPy, statsmodel, and scikit-learn, and Pandas UDFs allow data scientists not only to scale out their workloads, but also to leverage the Pandas APIs in Apache Spark.
The user-defined functions are executed by:
- Apache Arrow, to exchange data directly between JVM and Python driver/executors with near-zero (de)serialization cost.
- Pandas inside the function, to work with Pandas instances and APIs.
The Pandas UDFs work with Pandas APIs inside the function and Apache Arrow for exchanging data. It allows vectorized operations that can increase performance up to 100x, compared to row-at-a-time Python UDFs.
The example below shows a Pandas UDF to simply add one to each value, in which it is defined with the function called pandas_plus_one decorated by pandas_udf with the Pandas UDF type specified as PandasUDFType.SCALAR.
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf('double', PandasUDFType.SCALAR)
def pandas_plus_one(v):
# `v` is a pandas Series
return v.add(1) # outputs a pandas Series
spark.range(10).select(pandas_plus_one("id")).show()
The Python function takes and outputs a Pandas Series. You can perform a vectorized operation for adding one to each value by using the rich set of Pandas APIs within this function. (De)serialization is also automatically vectorized by leveraging Apache Arrow under the hood.
Python Type Hints
Python type hints were officially introduced in PEP 484 with Python 3.5. Type hinting is an official way to statically indicate the type of a value in Python. See the example below.
def greeting(name: str) -> str:
return 'Hello ' + name
The name: strindicates the name argument is of str type and the -> syntax indicates the greeting() function returns a string.
Python type hints bring two significant benefits to the PySpark and Pandas UDF context.
- It gives a clear definition of what the function is supposed to do, making it easier for users to understand the code. For example, unless it is documented, users cannot know if
greetingcan takeNoneor not if there is no type hint. It can avoid the need to document such subtle cases with a bunch of test cases and/or for users to test and figure out by themselves. - It can make it easier to perform static analysis. IDEs such as PyCharm and Visual Studio Code can leverage type annotations to provide code completion, show errors, and support better go-to-definition functionality.
Proliferation of Pandas UDF Types
Since the release of Apache Spark 2.3, a number of new Pandas UDFs have been implemented, making it difficult for users to learn about the new specifications and how to use them. For example, here are three Pandas UDFs that output virtually the same results:
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf('long', PandasUDFType.SCALAR)
def pandas_plus_one(v):
# `v` is a pandas Series
return v + 1 # outputs a pandas Series
spark.range(10).select(pandas_plus_one("id")).show()
from pyspark.sql.functions import pandas_udf, PandasUDFType
# New type of Pandas UDF in Spark 3.0.
@pandas_udf('long', PandasUDFType.SCALAR_ITER)
def pandas_plus_one(itr):
# `iterator` is an iterator of pandas Series.
return map(lambda v: v + 1, itr) # outputs an iterator of pandas Series.
spark.range(10).select(pandas_plus_one("id")).show()
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf("id long", PandasUDFType.GROUPED_MAP)
def pandas_plus_one(pdf):
# `pdf` is a pandas DataFrame
return pdf + 1 # outputs a pandas DataFrame
# `pandas_plus_one` can _only_ be used with `groupby(...).apply(...)`
spark.range(10).groupby('id').apply(pandas_plus_one).show()
Although each of these UDF types has a distinct purpose, several can be applicable. In this simple case, you could use any of the three. However, each of the Pandas UDFs expects different input and output types, and works in a different way with a distinct semantic and different performance. It confuses users about which one to use and learn, and how each works.
Furthermore, pandas_plus_one in the first and second cases can be used where the regular PySpark columns are used. Consider the argument of withColumn or the function with the combinations of other expressions such as pandas_plus_one("id") + 1. However, the last pandas_plus_one can only be used with groupby(...).apply(pandas_plus_one).
This level of complexity has triggered numerous discussions with Spark developers, and drove the effort to introduce the new Pandas APIs with Python type hints via an official proposal. The goal is to enable users to naturally express their pandas UDFs using Python type hints without confusion as in the problematic cases above. For example, the cases above can be written as below:
def pandas_plus_one(v: pd.Series) -> pd.Series:
return v + 1
def pandas_plus_one(itr: Iterator[pd.Series]) -> Iterator[pd.Series]:
return map(lambda v: v + 1, itr)
def pandas_plus_one(pdf: pd.DataFrame) -> pd.DataFrame:
return pdf + 1
New Pandas APIs with Python Type Hints
To address the complexity in the old Pandas UDFs, from Apache Spark 3.0 with Python 3.6 and above, Python type hints such as pandas.Series, pandas.DataFrame, Tuple, and Iterator can be used to express the new Pandas UDF types.
In addition, the old Pandas UDFs were split into two API categories: Pandas UDFs and Pandas Function APIs. Although they work internally in a similar way, there are distinct differences.
You can treat Pandas UDFs in the same way that you use other PySpark column instances. However, you cannot use the Pandas Function APIs with these column instances. Here are these two examples:
# Pandas UDF
import pandas as pd
from pyspark.sql.functions import pandas_udf, log2, col
@pandas_udf('long')
def pandas_plus_one(s: pd.Series) -> pd.Series:
return s + 1
# pandas_plus_one("id") is identically treated as _a SQL expression_ internally.
# Namely, you can combine with other columns, functions and expressions.
spark.range(10).select(
pandas_plus_one(col("id") - 1) + log2("id") + 1).show()
# Pandas Function API
from typing import Iterator
import pandas as pd
def pandas_plus_one(iterator: Iterator[pd.DataFrame]) -> Iterator[pd.DataFrame]:
return map(lambda v: v + 1, iterator)
# pandas_plus_one is just a regular Python function, and mapInPandas is
# logically treated as _a separate SQL query plan_ instead of a SQL expression.
# Therefore, direct interactions with other expressions are impossible.
spark.range(10).mapInPandas(pandas_plus_one, schema="id long").show()
Also, note that Pandas UDFs require Python type hints whereas the type hints in Pandas Function APIs are currently optional. Type hints are planned for Pandas Function APIs and may be required at some point in the future.
New Pandas UDFs
Instead of defining and specifying each Pandas UDF type manually, the new Pandas UDFs infer the Pandas UDF type from the given Python type hints at the Python function. There are currently four supported cases of the Python type hints in Pandas UDFs:
- Series to Series
- Iterator of Series to Iterator of Series
- Iterator of Multiple Series to Iterator of Series
- Series to Scalar (a single value)
Before we do a deep dive into each case, let’s look at three key points about working with the new Pandas UDFs.
- Although Python type hints are optional in the Python world in general, you must specify Python type hints for the input and output in order to use the new Pandas UDFs.
- Users can still use the old way by manually specifying the Pandas UDF type. However, using Python type hints is encouraged.
- The type hint should use
pandas.Seriesin all cases. However, there is one variant in whichpandas.DataFrameshould be used for its input or output type hint instead: when the input or output column is ofStructType.Take a look at the example below:
import pandas as pd from pyspark.sql.functions import pandas_udf df = spark.createDataFrame( [[1, "a string", ("a nested string",)]], "long_col long, string_col string, struct_col struct<col1:string>") @pandas_udf("col1 string, col2 long") def pandas_plus_len( s1: pd.Series, s2: pd.Series, pdf: pd.DataFrame) -> pd.DataFrame: # Regular columns are series and the struct column is a DataFrame. pdf['col2'] = s1 + s2.str.len() return pdf # the struct column expects a DataFrame to return df.select(pandas_plus_len("long_col", "string_col", "struct_col")).show()
Series to Series
Series to Series is mapped to scalar Pandas UDF introduced in Apache Spark 2.3. The type hints can be expressed as pandas.Series, ... -> pandas.Series. It expects the given function to take one or more pandas.Series and outputs one pandas.Series. The output length is expected to be the same as the input.
import pandas as pd
from pyspark.sql.functions import pandas_udf
@pandas_udf('long')
def pandas_plus_one(s: pd.Series) -> pd.Series:
return s + 1
spark.range(10).select(pandas_plus_one("id")).show()
The example above can be mapped to the old style with scalar Pandas UDF, as below.
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf('long', PandasUDFType.SCALAR)
def pandas_plus_one(v):
return v + 1
spark.range(10).select(pandas_plus_one("id")).show()
Iterator of Series to Iterator of Series
This is a new type of Pandas UDF coming in Apache Spark 3.0. It is a variant of Series to Series, and the type hints can be expressed as Iterator[pd.Series] -> Iterator[pd.Series]. The function takes and outputs an iterator of pandas.Series.
The length of the whole output must be the same length of the whole input. Therefore, it can prefetch the data from the input iterator as long as the lengths of entire input and output are the same. The given function should take a single column as input.
from typing import Iterator
import pandas as pd
from pyspark.sql.functions import pandas_udf
@pandas_udf('long')
def pandas_plus_one(iterator: Iterator[pd.Series]) -> Iterator[pd.Series]:
return map(lambda s: s + 1, iterator)
spark.range(10).select(pandas_plus_one("id")).show()
It is also useful when the UDF execution requires expensive initialization of some state. The pseudocode below illustrates the case.
@pandas_udf("long")
def calculate(iterator: Iterator[pd.Series]) -> Iterator[pd.Series]:
# Do some expensive initialization with a state
state = very_expensive_initialization()
for x in iterator:
# Use that state for the whole iterator.
yield calculate_with_state(x, state)
df.select(calculate("value")).show()
Iterator of Series to Iterator of Series can be also mapped to the old Pandas UDF style. See the example below.
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf('long', PandasUDFType.SCALAR_ITER)
def pandas_plus_one(iterator):
return map(lambda s: s + 1, iterator)
spark.range(10).select(pandas_plus_one("id")).show()
Iterator of Multiple Series to Iterator of Series
This type of Pandas UDF will be also introduced in Apache Spark 3.0, together with Iterator of Series to Iterator of Series. The type hints can be expressed as Iterator[Tuple[pandas.Series, ...]] -> Iterator[pandas.Series].
It has the similar characteristics and restrictions with Iterator of Series to Iterator of Series. The given function takes an iterator of a tuple of pandas.Series and outputs an iterator of pandas.Series. It is also useful when to use some states and when to prefetch the input data. The length of the entire output should also be the same as the length of the entire input. However, the given function should take multiple columns as input, unlike Iterator of Series to Iterator of Series.
from typing import Iterator, Tuple
import pandas as pd
from pyspark.sql.functions import pandas_udf
@pandas_udf("long")
def multiply_two(
iterator: Iterator[Tuple[pd.Series, pd.Series]]) -> Iterator[pd.Series]:
return (a * b for a, b in iterator)
spark.range(10).select(multiply_two("id", "id")).show()
This can also be mapped to the old Pandas UDF style as below.
from pyspark.sql.functions import pandas_udf, PandasUDFType
@pandas_udf('long', PandasUDFType.SCALAR_ITER)
def multiply_two(iterator):
return (a * b for a, b in iterator)
spark.range(10).select(multiply_two("id", "id")).show()
Series to Scalar
Series to Scalar is mapped to the grouped aggregate Pandas UDF introduced in Apache Spark 2.4. The type hints are expressed as pandas.Series, ... -> Any. The function takes one or more pandas.Series and outputs a primitive data type. The returned scalar can be either a Python primitive type, e.g., int, float, or a NumPy data type such as numpy.int64, numpy.float64, etc. Any should ideally be a specific scalar type accordingly.
import pandas as pd
from pyspark.sql.functions import pandas_udf
from pyspark.sql import Window
df = spark.createDataFrame(
[(1, 1.0), (1, 2.0), (2, 3.0), (2, 5.0), (2, 10.0)], ("id", "v"))
@pandas_udf("double")
def pandas_mean(v: pd.Series) -> float:
return v.sum()
df.select(pandas_mean(df['v'])).show()
df.groupby("id").agg(pandas_mean(df['v'])).show()
df.select(pandas_mean(df['v']).over(Window.partitionBy('id'))).show()
The example above can be converted to the example with the grouped aggregate Pandas UDF as you can see here:
import pandas as pd
from pyspark.sql.functions import pandas_udf, PandasUDFType
from pyspark.sql import Window
df = spark.createDataFrame(
[(1, 1.0), (1, 2.0), (2, 3.0), (2, 5.0), (2, 10.0)], ("id", "v"))
@pandas_udf("double", PandasUDFType.GROUPED_AGG)
def pandas_mean(v):
return v.sum()
df.select(pandas_mean(df['v'])).show()
df.groupby("id").agg(pandas_mean(df['v'])).show()
df.select(pandas_mean(df['v']).over(Window.partitionBy('id'))).show()
New Pandas Function APIs
This new category in Apache Spark 3.0 enables you to directly apply a Python native function, which takes and outputs Pandas instances against a PySpark DataFrame. Pandas Functions APIs supported in Apache Spark 3.0 are: grouped map, map, and co-grouped map.
Note that the grouped map Pandas UDF is now categorized as a group map Pandas Function API. As mentioned earlier, the Python type hints in Pandas Function APIs are optional currently.
Grouped Map
Grouped map in the Pandas Function API is applyInPandas at a grouped DataFrame, e.g., df.groupby(...). This is mapped to the grouped map Pandas UDF in the old Pandas UDF types. It maps each group to each pandas.DataFrame in the function. Note that it does not require for the output to be the same length of the input.
import pandas as pd
df = spark.createDataFrame(
[(1, 1.0), (1, 2.0), (2, 3.0), (2, 5.0), (2, 10.0)], ("id", "v"))
def subtract_mean(pdf: pd.DataFrame) -> pd.DataFrame:
v = pdf.v
return pdf.assign(v=v - v.mean())
df.groupby("id").applyInPandas(subtract_mean, schema=df.schema).show()
Grouped map type is mapped to grouped map Pandas UDF supported from Spark 2.3, as below:
import pandas as pd
from pyspark.sql.functions import pandas_udf, PandasUDFType
df = spark.createDataFrame(
[(1, 1.0), (1, 2.0), (2, 3.0), (2, 5.0), (2, 10.0)], ("id", "v"))
@pandas_udf(df.schema, PandasUDFType.GROUPED_MAP)
def subtract_mean(pdf):
v = pdf.v
return pdf.assign(v=v - v.mean())
df.groupby("id").apply(subtract_mean).show()
Map
Map Pandas Function API is mapInPandas in a DataFrame. It is new in Apache Spark 3.0. It maps every batch in each partition and transforms each. The function takes an iterator of pandas.DataFrame and outputs an iterator of pandas.DataFrame. The output length does not need to match the input size.
from typing import Iterator
import pandas as pd
df = spark.createDataFrame([(1, 21), (2, 30)], ("id", "age"))
def pandas_filter(iterator: Iterator[pd.DataFrame]) -> Iterator[pd.DataFrame]:
for pdf in iterator:
yield pdf[pdf.id == 1]
df.mapInPandas(pandas_filter, schema=df.schema).show()
Co-grouped Map
Co-grouped map, applyInPandas in a co-grouped DataFrame such as df.groupby(...).cogroup(df.groupby(...)), will also be introduced in Apache Spark 3.0. Similar to the grouped map, it maps each group to each pandas.DataFrame in the function but it groups with another DataFrame by common key(s) and then the function is applied to each cogroup. Likewise, there is no restriction on the output length.
import pandas as pd
df1 = spark.createDataFrame(
[(1201, 1, 1.0), (1201, 2, 2.0), (1202, 1, 3.0), (1202, 2, 4.0)],
("time", "id", "v1"))
df2 = spark.createDataFrame(
[(1201, 1, "x"), (1201, 2, "y")], ("time", "id", "v2"))
def asof_join(left: pd.DataFrame, right: pd.DataFrame) -> pd.DataFrame:
return pd.merge_asof(left, right, on="time", by="id")
df1.groupby("id").cogroup(
df2.groupby("id")
).applyInPandas(asof_join, "time int, id int, v1 double, v2 string").show()
Conclusion and Future Work
The upcoming release of Apache Spark 3.0 (read our preview blog for details). will offer Python type hints to make it simpler for users to express Pandas UDFs and Pandas Function APIs. In the future, we should consider adding support for other type hint combinations in both Pandas UDFs and Pandas Function APIs. Currently, the supported cases are only few of many possible combinations of Python type hints. There are also other ongoing discussions in the Apache Spark community. Visit Side Discussions and Future Improvement to learn more.
Try out these new capabilities today for free on Databricks as part of the Databricks Runtime 7.0 Beta.
--
Try Databricks for free. Get started today.
The post New Pandas UDFs and Python Type Hints in the Upcoming Release of Apache Spark 3.0™ appeared first on Databricks.