Friday, 5 April 2019

‘We Should Be a Lot Further Along Than We Are’ Says Former DOD No. 2 of AI

The Pentagon has recently made a huge push into adopting operationalized artificial intelligence for military applications. But according to former Deputy Secretary of Defense Bob Work, the U.S. military is in grave danger of falling behind the Chinese in the race to develop AI and losing its competitive advantage if it doesn’t go all-in.

“I can’t shake the nagging feeling … that we should be a lot further along than we are and we’re losing ground to our competitors,” said Work, now a distinguished senior fellow for defense and national security at the bipartisan Center for a New American Security.

Responsible for the Third Offset Strategy during his time as deputy secretary of Defense, Work helped establish the Pentagon’s algorithmic warfare efforts that now serve as a precursor for the Joint AI Center, launched last summer to lead the military’s AI efforts. Earlier this year, he was named as co-chair of the National Security Commission for Artificial Intelligence.

“If we’re going to succeed against a competitor like China that is all-in in this competition — I mean they are all in, from the top leadership down to the commanders in the field — we’re going to have to grasp the inevitability of AI and adapt our own innovation culture and behavior so that AI has a chance to take hold,” he said Wednesday at AFCEA’s Artificial Intelligence and Machine Learning Summit.

Specifically, Work said the DOD needs “small plays” going on department wide — “substantial, sustained experimentation using these technologies, widespread applications being applied by the services across all operating domains.”

“We will not be able, in my view, to defeat China in this competition, unless we change the way we’re going after this in a broad way,” he said.

In a large sense, that’s the point of the JAIC: to help DOD from, the departmental view, wrap its arms around the hundreds of AI projects ongoing around the military. Lt. Gen. Jack Shanahan, who heads JAIC, agreed Wednesday that “we have to move faster, to do better, that’s what we’re really trying to do right now.”

“If we project 20 years into the future, and we’re on the cusp of a major conflict with a peer competitor, if at that point we have a truly AI-enabled DOD force, that by itself will not imply that we will win the conflict,” Shanahan said in a separate keynote. “If we don’t have a fully AI-enabled force, we will incur an unacceptably high risk of losing. That’s how important this is to our national security.”

Talking ethics

As they often do, the conversations Wednesday on AI found their way back to the ethical implication of the technology being used in connection to lethality.

Work and Shanahan both agreed that it’s healthy to have a dialogue about ethics — but they also pointed to misconceptions about the U.S. military’s ethical use of technology in general.

“I would argue that the United States military is the most ethical military force in the history of warfare, and we think the shift to AI-enabled weapons will continue this trend,” Work said. The existing policy, which predates any of this current work, he said, “is very clear that these weapons have to be consistent with the laws of armed conflict, supporting the principles of distinction and proportionality, and it has been DOD policy since 2012, three years before the Third Offset, that every weapon we field must be designed and deployed to allow commanders and operators to exercise appropriate levels of human judgment in the lethal application of force.”

Shanahan explained there “are grave misperceptions about what DOD is actually working on” with AI.

“In my experience, in the last two years, what I’ve found is there’s the assumption in some corners that the DOD in a back laboratory somewhere in a basement of a building has got a free-will AGI, artificial general intelligence, that’s going to roam indiscriminately across the battlefield,” he said. “We do not.”

Instead, he explained, DOD is looking to adopt applications of artificial narrow intelligence — “it’s for specific problems, and just like every other technology we ever work with in the department, from the beginning we take into this question of what is the technology meant to be used for? What are the ethical, safety and law implications of using that technology?”

Read the source article in fedscoop.

Machine Learning Ultra-Brittleness and Object Orientation Poses: The Case of AI Self-Driving Cars

By Lance Eliot, the AI Trends Insider

Take an object nearby you and turn it upside down. Please don’t do this to something or someone that would get upset at your suddenly turning them upside down. Assuming that you’ve turned an object upside down, look at it. Do you still know what the object is? I’d bet that you do.

But why would you? If you were used to seeing it right-side up, presumably you should be baffled at what the object is, now that you’ve turned it upside down. It doesn’t look like it did a moment ago. A moment ago, the bottom was, well, on the bottom. The top was on the top. Now, the bottom is on the top, and the top is on the bottom. I dare say that you should be completely puzzled about the object. It is unrecognizable now that it has been flipped over.

I’m guessing that you are puzzled that I would even suggest that you should be puzzled. Of course, you recognize what the object is. No big deal. It seems silly perhaps to assert that the mere act of turning the object upside down should impact your ability to recognize the object. You might insist that the object is still the same object that it was a moment ago. No change has occurred. It is simply reoriented.

Not so fast. Your ability as a grown adult is helping you quite a bit on this seemingly innocuous task. For you, it has been years upon years of cognitive maturation that makes things so easy to perceive an object when reoriented.

I could get you to falter somewhat by showing you an object that I had hidden behind my back and suddenly showed it to you, only showing it to you while it is being held upside down. Without first my showing it to you in a right-side up posture, the odds are that it would take you a few moments to figure out what the upside-down object was.

A Father’s Story About Reorienting Objects

This discussion causes me to hark back to when my daughter was quite young.

She had a favorite doll that had a big grin on the doll, permanently in place. There were a series of small buttons that shaped the mouth and it was curved in a manner that made it look like a smiley face type of grin. When you looked at the doll, you almost instinctively would react to the wide smile and it would spark you to smile too. She really liked the doll, it was her favorite toy and not to be trifled with.

One day, we were sitting at the dinner table and I opted to turn the doll upside down. I asked my daughter whether the doll was smiling or whether the doll was frowning. Though I realize you cannot at this moment see the doll that I am referring to, I’m sure you realize that the doll was still smiling, but when the doll was turned upside down, the smile would be upside down and resemble a frown.

My daughter said that the doll was sad, it was frowning.

I turned the doll right-side up.

And now what is the expression, I asked my daughter.

The doll is smiling again, she said.

I explained that the doll had always been smiling, even when turned upside down. This got a look of puzzlement on my daughter’s face. By the way, I was potentially on the verge of trifling with her favored toy, so I assure you that I carried out this activity with great respect and care.

She challenged me to turn the doll upside down again. I did so.

My daughter stood-up, and tried to do a handstand, flipping herself upside down. Upon quasi-doing so, she gazed at the doll, and could see that the doll was still smiling. She agreed that the doll was still smiling and retracted her earlier indication that it had been frowning.

I waited about a week and tried to pull this stunt again. This time, she responded instantly that the doll was still smiling, even after I had turned it upside down. She obviously had caught on.

When I tried this once again about two weeks later, she said the doll was sad. This surprised me and I wondered if she had perchance forgotten our earlier stints. When I asked her why the doll was sad, she told me it was because I keep turning her upside down and she’s not fond of my doing so. Ha! I was put into my place.

The overall point herein is that when we are very young, being able to discern objects that are upside down can be quite difficult. You’ve not yet modeled in your mind the notion of reorienting objects and being able to rotate them, doing so to then recognize them as readily. Sure, my daughter knew that the doll was still the doll, but the smile that was a frown suggested an as-yet developed sense of reorienting of objects.

Human Mental Capabilities in Reorienting Objects

What makes our learning capabilities so impressive is that you don’t just fixate on a particular object and instead you ultimately generalize to objects all told. My daughter was able to not only figure out the doll once it was upside down, she generalized this modeling to be able to then figure out other objects that were turned upside down. If I showed her an object in the right-side up position first, it was usually relatively easy for her to comprehend the object once I had turned it upside down.

Turning an object upside down, prior to presenting it, can be a bit of a challenge to your identifying an object, when presented to someone, even for adults. We are momentarily caught off-guard by the untoward orientation (assuming that you don’t normally see it upside down).

Your mind tries to examine the upside-down object and perhaps reorients the object in your mind, creating a picture in your mind, and flipping the picture to a right-side up orientation to make sense of it. You then match the reoriented mental image of the object to your stored right-side up images, and voila, you identify what the real-world object is.

Or, it could be that the mind takes a known right-side up image that’s already in its stored memory, and for which you believe it might be, and flips over the stored image that’s in your head, and then matches it to the object that you are seeing positioned as upside down, trying to decide if it is indeed that object. That’s another plausible way to do this.

There have been lots of cognitive and psychological experiments trying to figure out the mental mechanisms in the brain that aid us when dealing with the reorientation of objects. Theories abound about how are brain actually figures these things out. I’ve so far suggested or implied that we keep an image of objects in our mind. Like a picture. But, that’s hard to prove.

Maybe it is some kind of calculus in our minds and there isn’t an object image per se being used. It could be a bunch of formulas. Maybe our minds are a vast collection of geometric formulas. Or, it could be a bunch of numbers. Perhaps our minds turn everything into a kind of binary code and there aren’t any images per se in our minds (well, I suppose it could be an image represented in a binary code).

The actual brain functioning is still a mystery and other than seemingly considerable and at times clever experiments, we cannot say for absolute certainty how the brain does this for us. Efforts in neuroscience continue to push forward, trying to nail down the mechanical biological and chemical plumbing of the brain.

Range of Reorienting Objects and Their Poses

I’ve focused on the idea of completely turning an object upside down. That’s not the only way to confuse our minds about an object.

You can turn an object on its side, which might also make things hard for you to then recognize the object. Usually, we quickly guess at the object when it is only partially reoriented and can seemingly do a pretty good guess at what it is. Turning the object upside down seems to be a more extreme variant, nonetheless even some milder reorientation can still cause us to pause or maybe even misclassify the object.

If I were to take an object and slowly rotate it, the odds are that you would be able to accurately say what the object is, assuming you watched it during the rotations. When I suddenly show you an object that has already been somewhat rotated, you have no initial basis to use as an anchor, and therefore it is more challenging to figure out what the object might be.

Familiarity plays a big part of this too. If I did a series of rotations of the object, and you were staring at it, your mind seems to be able to get used to those orientations. Thus, if later on, I suddenly spring upon you that same object in a rotated posture, you are more apt to quickly know what it is, due to having seen it earlier in the rotated position.

In that sense, I can essentially train your mind about what an object looks like in a variety of orientations, making it much easier for you to later on recognize it, when it is in one of those orientations. Maybe your mind has taken snapshots of each orientation. Or, maybe your mind is able to apply some kind of mental algorithm to the orientations of that objects. Don’t know.

People that deal with a multitude of orientations of objects tend to get better and better at the object reorientation task. I used to work for a CEO that had a trick plane. He would take me up in it, usually on our lunch break at work (our office was nearby an airport). He would do barrel rolls and all kinds of tricky flight maneuvers. I learned right away to not eat lunch before we went on these flights (think about the vaunted “vomit comet”).

In any case, he was able to “see” the world around us quite well, in spite of the times when we were flying upside down. For me, the world was quite confusing looking when we were upside down. I had a difficult time with it. Then again, I’ve never been the type to enjoy those roller coaster rides that turn you upside down and try to scare the heck out of you.

AI Self-Driving Cars and Object Orientations in Street Scenes

What does this have to do with AI self-driving cars?

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. One of the major concerns that we have, and the auto makers have, and tech firms have, pertains to Machine Learning or Deep Learning that we are all using today, and which tends to be ultra-brittle when it comes to objects that are reoriented.

This is bad because it means that the AI system might either not recognize an object due to the orientation of it, or the AI might misclassify an object, and end-up tragically getting the self-driving car into a precarious situation because of it.

Allow me to elaborate.

I’d like to first clarify and introduce the notion that there are varying levels of AI self-driving cars. The topmost level is considered Level 5. A Level 5 self-driving car is one that is being driven by the AI and there is no human driver involved. For the design of Level 5 self-driving cars, the auto makers are even removing the gas pedal, brake pedal, and steering wheel, since those are contraptions used by human drivers. The Level 5 self-driving car is not being driven by a human and nor is there an expectation that a human driver will be present in the self-driving car. It’s all on the shoulders of the AI to drive the car.

For self-driving cars less than a Level 5, there must be a human driver present in the car. The human driver is currently considered the responsible party for the acts of the car. The AI and the human driver are co-sharing the driving task. In spite of this co-sharing, the human is supposed to remain fully immersed into the driving task and be ready at all times to perform the driving task. I’ve repeatedly warned about the dangers of this co-sharing arrangement and predicted it will produce many untoward results.

For my overall framework about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

For the levels of self-driving cars, see my article: https://aitrends.com/selfdrivingcars/richter-scale-levels-self-driving-cars/

For why AI Level 5 self-driving cars are like a moonshot, see my article: https://aitrends.com/selfdrivingcars/self-driving-car-mother-ai-projects-moonshot/

For the dangers of co-sharing the driving task, see my article: https://aitrends.com/selfdrivingcars/human-back-up-drivers-for-ai-self-driving-cars/

Let’s focus herein on the true Level 5 self-driving car. Much of the comments apply to the less than Level 5 self-driving cars too, but the fully autonomous AI self-driving car will receive the most attention in this discussion.

Here’s the usual steps involved in the AI driving task:

  •         Sensor data collection and interpretation
  •         Sensor fusion
  •         Virtual world model updating
  •         AI action planning
  •         Car controls command issuance

Another key aspect of AI self-driving cars is that they will be driving on our roadways in the midst of human driven cars too. There are some pundits of AI self-driving cars that continually refer to a utopian world in which there are only AI self-driving cars on the public roads. Currently there are about 250+ million conventional cars in the United States alone, and those cars are not going to magically disappear or become true Level 5 AI self-driving cars overnight.

Indeed, the use of human driven cars will last for many years, likely many decades, and the advent of AI self-driving cars will occur while there are still human driven cars on the roads. This is a crucial point since this means that the AI of self-driving cars needs to be able to contend with not just other AI self-driving cars, but also contend with human driven cars. It is easy to envision a simplistic and rather unrealistic world in which all AI self-driving cars are politely interacting with each other and being civil about roadway interactions. That’s not what is going to be happening for the foreseeable future. AI self-driving cars and human driven cars will need to be able to cope with each other.

For my article about the grand convergence that has led us to this moment in time, see: https://aitrends.com/selfdrivingcars/grand-convergence-explains-rise-self-driving-cars/

See my article about the ethical dilemmas facing AI self-driving cars: https://aitrends.com/selfdrivingcars/ethically-ambiguous-self-driving-cars/

For potential regulations about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/assessing-federal-regulations-self-driving-cars-house-bill-passed/

For my predictions about AI self-driving cars for the 2020s, 2030s, and 2040s, see my article: https://aitrends.com/selfdrivingcars/gen-z-and-the-fate-of-ai-self-driving-cars/

Artificial Neural Networks (ANN) and Deep Neural Networks (DNN)

Returning to the topic of object orientation, let’s consider how today’s Machine Learning and Deep Learning works, along with why it is considered at times to be ultra-brittle. We’ll also mull over how this ultra-brittleness can spell sour outcomes for the emerging AI self-driving cars.

Take a look at Figure 1.

Suppose I decide to craft an Artificial Neural Network (ANN) that will aid in finding street signs, cars, and pedestrians inside of images or video streaming of a camera that is on a self-driving car. Typically, I would start by finding a large dataset of traffic setting images that I could use to train my ANN. We want this ANN to be as full-bodied as we can make it, so we’ll have a multitude of layers and compose it of a large number of artificial neurons, thus we might refer to this kind of more robust ANN as a Deep Neural Network (DNN).

Datasets Essential to Deep Learning

You might wonder how I will come upon the thousands upon thousands of images of traffic scenes. I need a rather large set of images to be able to appropriately train the DNN. I don’t want to have to go outside and start taking pictures, since it would take me a long time to do so and be costly to upload them all and store them. My best bet would be to go ahead and use datasets that already exist.

Indeed, some would say that the reason we’ve seen such great progress in the application of Deep Learning and Machine Learning is because of the efforts by others to create large-scale datasets that we can all use to do our training of the ANN or DNN. We stand on the shoulders of those that went to the trouble to put together those datasets, thanks.

This also though means there is a kind of potential vulnerability that is taking place, one that is not so obvious. If we all use the same datasets, and if those datasets have particular nuances in them, it means that we all are also going to be having a similar impact on our ANN and DNN trainings. I’ll in a moment provide you with an example involving military equipment images that can highlight this vulnerability.

For some well-known ANN/DNN training datasets, see my article: https://www.aitrends.com/selfdrivingcars/machine-learning-benchmarks-and-ai-self-driving-cars/

For my article about Deep Learning and plasticity, see: https://www.aitrends.com/selfdrivingcars/plasticity-in-deep-learning-dynamic-adaptations-for-ai-self-driving-cars/

For the vaunted one-shot Deep Learning goal, see my article: https://www.aitrends.com/selfdrivingcars/seeking-one-shot-machine-learning-the-case-of-ai-self-driving-cars/

For my article about the use of compressive sensing, see: https://www.aitrends.com/selfdrivingcars/compressive-sensing-ai-self-driving-cars/

Once I’ve got my dataset or datasets and readied my DNN to be trained, I would run the DNN over and over, trying to get it to find patterns in the images.

I might do so in a supervisory way, wherein I provide an indication of what I want it to find, such as I might give the DNN guidance toward discovering images of school buses or maybe of fire trucks or perhaps scooters. It might be that I opt to do this in an unsupervised fashion and allow the DNN to find whatever it finds, and provide an indication of what those objects that it clusters or classifies are. For example, those yellow lengthy blobs that have big tires and lots of windows are school buses.

I am hoping that the DNN is generalizing sufficiently about the objects, in the sense that if a yellow school bus is bright yellow, it is still a school bus, while if it is maybe a dull yellow due to faded paint and dirt and grime, the DNN should still be classifying it into the school bus category. I mention this because you typically do not have an easy way to make the DNN explain what it is using to find and classify the objects within the image. Instead, you merely hope and assume that if it seems to be able to find those yellow buses, it presumably is using useful criteria to do so.

There is the famous story that highlights the dangers of making this kind of assumption about the manner in which the pattern matching is taking place. The story goes that there were pictures of Russia military equipment, like tanks and cannons, and there were pictures of United States military equipment. Thousands of images that had those kinds of equipment were fed into an ANN. The ANN seemed to be able to discern between the Russian military equipment and the United States military equipment, of which we would presume that it was due to the differences in the shape and designs of their respective tanks and cannons.

Turns out that upon further inspection, the pictures of the Russian military equipment were all grainy and slightly out of focus, while the United States military equipment pictures were crisp and bright. The ANN pattern matched on the background and lighting aspects, rather than the shape of the military equipment itself. This was not readily discerned at first because the same set of images were used to train the ANN and then test it. Thus, the test set were also grainy for the Russian equipment and crisp for the U.S. equipment, misleading one into believing that the ANN was doing a generalized job of gauging the object differences, when it was not doing so.

This highlights an important aspect for those using Machine Learning and Deep Learning, namely trying to ferret out how your ANN or DNN is achieving its pattern matching. If you treat it utterly like a black box, there might be ways in which the pattern matching has landed that won’t be satisfactory for use when the ANN or DNN is used in real-world ways. You might have thought that you did a great job, but once the ANN or DNN is exposed to other images, beyond your datasets, it could be that the characteristics used to classify objects is revealed as brittle and not what you had hoped for.

Considering Deep Learning as Brittle and Ultra-Brittle

By the word “brittle” I am referring to the notion that the ANN or DNN is not doing a full-bodied kind of pattern matching and will therefore falter or fall-down on doing what you presumably want it to do. In the case of the tanks and cannons, you likely wanted the patterns to be about the shape of the tank, its turret, its muzzle, its treads, etc. Instead, the pattern matching was about the graininess of the images. That’s not going to do much good when you try to use the ANN or DNN in a real-world environment to detect whether there is a Russian tank or a United States tank ahead of you.

Let’s liken this to my point about the yellow school bus. If the ANN or DNN is pattern matching on the color of yellow, and if perchance all of most of the images in my dataset were of bright yellow school buses, it could be that the matching is being done by that bright yellow color. This means that if I think that my ANN or DNN is good to go, and it encounters a school bus that is old, faded in yellow color, and perhaps covered with grime, the ANN or DNN might declare that the object is not a school bus. A human would tell you it was a school bus, since the human presumably is looking at a variety of characteristics, including the wheels, the shape of the bus, the windows, and the color of the bus.

One of the ways in which the brittleness of the ANN or DNN can be exploited involves making use of adversarial images. The notion is to confuse or mislead the trained ANN or DNN into misclassifying an object. This might be done by a bad actor, someone hoping to cause the ANN or DNN to falter. They can take an image, make some changes to it, feed it into the ANN or DNN that you’ve crafted, and potentially get the ANN or DNN to say that the object is something other than what it is.

Perhaps one of the more famous examples of this kind of adversarial trickery involves the turtle image that an ANN or DNN was fooled into believing was actually an image of a gun. This can be done by making changes in the image of the turtle. Those changes are enough to have the ANN or DNN no longer pattern match it to being a frog and instead pattern match it to being a gun. What makes these adversarial attacks so alarming is that the turtle might still look like a turtle to the human eye, and the changes made to fool the ANN or DNN are at the pixel level, being so small that the human eye doesn’t readily see the difference.

One of the more startling examples of this adversarial trickery involved a one-pixel change that caused an apparent image of a dog to be classified by a DNN as a cat, which goes to show how potentially brittle these systems can be. Those that study these kinds of attacks will often used “differential evolution” or DE to try and find the least amount of change that is the least apparent to humans, aiming to then fool the ANN or DNN and yet make it very hard for a human eye to realize what has been done.

These changes to images are also often referred to as adversarial perturbations.

Remember that I earlier said that by using the same datasets we are somewhat vulnerable, well, a bad actor can study those datasets too, and try to find ways to undermine or undercut an ANN or DNN that has been trained via the use of those datasets. The dataset giveth and it taketh, one might say. By having large-scale datasets readily available, it means the good actors can more readily develop their ANN or DNN, but it also means that the bad actors can try to figure out ways to subvert those good guy ANN and DNN’s, doing so by discovering devious adversarial perturbations.

Not all adversarial perturbations need to be conniving, and we might use the adversarial gambit for good purposes too. When you are testing the outcome of your ANN or DNN training, it would be wise to try and do some adversarial perturbations to see what you can find, meaning that you are trying to use this technique to detect your own brittleness. By doing so, hopefully you will be able to then try to shore-up that brittleness. Might as well use the attack for the purposes of discovery and mitigation.

For my article about security aspects of AI self-driving cars, see: https://www.aitrends.com/ai-insider/ai-deep-learning-backdoor-security-holes-self-driving-cars-detection-prevention/

For AI brittleness aspects, see my article: https://www.aitrends.com/ai-insider/goto-fail-and-ai-brittleness-the-case-of-ai-self-driving-cars/

For ensemble Machine Learning, see my article: https://www.aitrends.com/selfdrivingcars/ensemble-machine-learning-for-ai-self-driving-cars/

For my article about Federated Machine Learning, see: https://www.aitrends.com/selfdrivingcars/federated-machine-learning-for-ai-self-driving-cars/

For explanation-AI, see my article: https://www.aitrends.com/selfdrivingcars/explanation-ai-machine-learning-for-ai-self-driving-cars/

I’ve so far offered the notion that the images might differ by the color of the object, such as the variants of yellow for a school bus. A school bus also has a number of wheels and tires, which are somewhat large, relative to smaller vehicles. A scooter only has two wheels and those tires are quite a bit smaller than a buses tire.

Imagine looking at image after image of school buses and trying to figure out what features allow them to formulate in your mind that they are school buses. You want to discover a wide enough set of criteria that it is not going to be brittle, yet you also don’t want to be overly broad that you then start classifying say trucks as school buses simply because both of those types of transport have larger tires.

Let’s add a twist to this. I told you about my daughter and her doll, involving my flipping the doll upside down and asking my daughter whether she could discern if the doll was smiling or frowning. I was not changing any of the actual features of the doll. It was still the same doll, but it was reoriented. That’s the only change I made.

Suppose we trained an ANN or DNN with thousands upon thousands of images of yellow buses. The odds are that the pictures of these yellow buses are primarily all of the same overall orientation, namely driving along on a flat road or maybe in a parked spot, sitting perfectly upright. The bus is right-side up.

You probably would assume that the ANN or DNN is pattern matching in such a manner that it doesn’t matter the nature of the orientation of the bus. You would take it for granted that the ANN or DNN “must” realize that the orientation doesn’t matter, a bus is still a bus, regardless at what angle and too when upside down.

If we were to tilt the bus, a human would likely still be able to tell you that it is a school bus. I could probably turn the bus completely upside down, if I could do so, and you’d still be able to discern that it is a school bus. I remember one day I was driving along and drove past an accident scene involving a car that had completely flipped upside down and was sitting at the side of the road. I marveled at seeing a car that was upside down. Notice that I could instantly detect it was a car, there was no confusion in my mind that it was anything other than a car, in spite of the fact that it was entirely upside down.

Fascinating Study of Poses Problems in Machine Learning

A fascinating new study by researchers at Auburn University and Adobe provides a handy warning that orientation should not be taken for granted when training your Deep Learning or Machine Learning system. Researchers Michael Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, and Anh Nguyen investigated the vulnerability of DNN’s, doing so by using adversarial techniques, primarily involving rotating or reorienting objects in images. These mainly were DNN’s that had been trained on rather popular datasets, such as ImageNet and MS COCO. Their study can be found here: https://arxiv.org/pdf/1811.11553.pdf

One aspect about the rotation or reorienting of an object that you might have noticed herein is that I’ve been suggesting that the objects are 2D and you are merely tilting them or putting them upside down. Given that for most real-world objects like school buses and cars, they are 3D objects, you can do the rotations or reorienting in three dimensions, altering the yaw, pitch, and roll of the object.

In the research study done at Auburn University and with Adobe, the researchers opted to try coming up with rather convincing looking Adversarial Examples (AX), contriving them to be Outside-of-the-Distribution (OoD) of the training datasets.

For example, using a Photoshop-like technique, they took an image of a yellow school bus and tilted it a few degrees, and went further to conjure up an image of the bus turned on its side. These were made to look like the images in the datasets, including having a background of a road that other school buses in the dataset were also shown upon. This helps to make these adversarial perturbations be focused more so on the object of interest, the school bus in this case, and not have the ANN or DNN hopefully be getting distracted by the background of the image as a giveaway.

To the human eye, these adversarial changes are blatantly obvious.

There wasn’t an effort to hide the perturbations by infusing them at a pixel level. You can look at a picture and immediately discern that there is a school bus in the picture, though you might certainly wonder why the school bus is at a tilt. It wasn’t bizarre pe se in that some of the reoriented images were plausible. A yellow school bus laying on its side, on a road, well, it could have gotten into an accident and ended-up in that position.

Some of the images might be questioned, like a fire truck that seems to be flying in the air, but I would also bet that if you had a fire truck that went off a bridge or a ramp, you’d be able to get the same kind of reorientation.

For the school bus, some of the reorientations caused the ANN or DNN to report that it was a garbage truck, or that it was a punching bag, or that it was a snowplow. The punching bag classification seems to make sense in that the yellow bus was dangling as though it was being held by its tailpipe, and since it is yellow, it might seem characteristic of a yellow punching bag that is hanging from a ceiling and ready to be punched. I don’t know for sure that this is the criteria used by the ANN or DNN, but it seems like a reasonable guess as based on the misclassification.

Of the objects that they decided to convert from their normal or canonical poses in the images, and reoriented to a different pose stance, they were able to get the selected DNN’s to do a misclassification 97% of the time. You might assume that this only happens when the pose is radically altered. You’d be wrong. They tried various pose changes and seemed to find that with just an approximate 10% yaw change, an 8% pitch change, or a 9% roll change, it was enough to fool the DNN.

You might also be thinking that this reorientation only causes a misclassification about school buses, and maybe it doesn’t apply to other kinds of objects. Objects they studied included a school bus, park bench, bald eagle, beach wagon, tiger cat, German shepherd, motor scooter, jean, street sign, moving van, umbrella, police van, and a trailer truck.

That’s enough of a variety that I think we can reasonably suggest that it showcases a diversity of objects and therefore is generalizable as a potential concern.

Variant Poses Suggest Ultra-Brittleness

Many people refer to today’s Machine Learning and Deep Learning as brittle. I’ll go even further and claim that it is ultra-brittle. I do so to emphasize the dangers we face by today’s ANN and DNN applications. Not only are they brittle with respect to the feature’s detection of objects, such as a bright yellow versus a faded yellow, they are brittle when you simply rotate or reorient an object. That’s why I am going to call this as being ultra-brittle.

If today’s ANN and DNN could deal with most rotations and reorientations, being able to still do a decent job of classifying objects, and they were only confounded by extraordinary poses, I would likely backdown from saying they were ultra-brittle and settle on brittle. The part that catches your attention and your throat is that it doesn’t take much of a perturbation to get the usual ANN or DNN to do a misclassification.

In the real-world, when an AI self-driving car is zooming along at 80 miles per hour, you certainly don’t want the on-board AI and ANN or DNN to misclassify objects due to their orientation.

I remember one harrowing time that I was driving my car and another car, going in the opposing direction, came across a tilted median that was intended to protect the separated directions of traffic. The car was on an upper street and I was on a lower street.

I don’t know whether the driver was drunk or maybe had fallen asleep, but in any case, he dove down toward the lower street. His car was at quite an obtuse angle.

What would an AI self-driving car have determined? Suppose the sensors detected the object but somehow gave it a more harmless classification such as naming it a wild animal, or a tumble weed? If so, the AI action planner might decide that there is no overt threat and not try to guide the self-driving car away, and instead assume that it would be safer to proceed ahead and ram the object, like ramming a deer that has suddenly appeared in the roadway.

I realize that some might shirk off the orientation aspects by suggesting that you are rarely going to see a school bus at an odd angle, or a fire truck, or anything else. I’m not so convinced. If we tried to come up with examples of reoriented objects in real-world settings, I’m betting we could readily identify numerous realistic situations. And, if we are going to ultimately have millions of AI self-driving cars on the roadways, the odds of that many self-driving cars eventually encountering “odd” poses is going to be relatively high.

For my article about safety and AI, see: https://www.aitrends.com/selfdrivingcars/safety-and-ai-self-driving-cars-world-safety-summit-on-autonomous-tech/

For the Boeing 737 situation, see my article: https://www.aitrends.com/ai-insider/boeing-737-max-8-and-lessons-for-ai-the-case-of-ai-self-driving-cars/

For the linear non-threshold aspects, see: https://www.aitrends.com/ai-insider/linear-no-threshold-lnt-and-the-lives-saved-lost-debate-of-ai-self-driving-cars/

For my article containing my Top 10 predications about AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/top-10-ai-trends-insider-predictions-about-ai-and-ai-self-driving-cars-for-2019/

What To Do About the Poses Problem

There are several ways we can gradually deal with this issue of the poses problem.

They include:

  •         Improve the ANN or DNN algorithms being used
  •         Increasing the scale of the ANN or DNN used
  •         Ensure that the datasets include variant poses
  •         Use adversarial techniques to ferret out and then mitigate poses issues
  •         Improve ANN or DNN explanatory capabilities
  •         Other

Researchers should be trying to devise Deep Learning and Machine Learning algorithms that can semi-automatically try to cope with the poses problem. This might involve the ANN or DNN itself opting to rotate or reorient objects, even though the reorientation wasn’t fed into the ANN or DNN via the training dataset. You might liken this to how humans in our minds seem to be able to do rotations of objects, even though we might not have the object in front of us in a rotated position.

If you think that the solution should focus more so on the dataset rather than the ANN or DNN itself, presumably we can try to include more variants of poses of objects into a dataset.

This is not straightforward, unfortunately.

It seems fair to assume that you are not likely to get actual pictures of those objects in a variety of orientations, naturally, and so you’d have to synthesize it. The synthesis itself will need to be convincing, else the images will be tagged by the ANN or DNN simply due to some other factor, akin to my example earlier about the grainy nature of the military equipment images.

Also, keep in mind that you need enough of the reoriented object images to make a difference when the ANN or DNN is doing the training on the dataset. If you have a million pictures of a school bus in a right-side up pose and have a handful of the bus in a tilted posture, the odds are that the pattern matching is going to overlook or ignore or cast aside as noise the tilted postures. This takes us back to the one-shot learning problem too.

You could be tempted to suggest that the dataset maybe should have many of the tilted poses, perhaps more so than the number of poses in a right-side up position. Well, this could be undesirable too. The pattern matching might become reliant on the tilted postures and not be able to recognize sufficiently when the object is in its normal or canonical position.

Darned if you do, darned if you don’t.

The Key 4 A’s of Datasets for Deep Learning

When we put together our datasets, we tend to think of the mixture in the following way:

  •         Anticipated poses
  •         Adaptation poses
  •         Aberration poses
  •         Adversarial poses

It’s the 4 A’s of poses or orientations.

We want to have some portion of the dataset with the anticipated poses, which are usually the right-side up or canonical orientations.

We want to have some portion of the dataset with the adaptation poses, namely postures that you could reasonably expect to occur from time-to-time in the real-world. It’s not the norm, but nor is it something that is extraordinary or unheard of in terms of orientation.

We want to ensure that there are a sufficient number of aberrations poses, entailing orientations that are quite rare and seemingly unlikely.

And we want to have some inclusion of adversarial poses that are let’s say concocted and would not seem to ever happen naturally, but for which we want to use so that if someone is determined to attack the ANN or DNN, it has already encountered those orientations. Note this is not the pixel-level kind of attacks preparation, which is handled in other ways.

You need to be reviewing your datasets to ascertain what mix you have of the 4 A’s. Is it appropriate for what you are trying to achieve with your ANN or DNN? Does the ANN or DNN have enough sensitivity to pick-up on the variants? And so on.

Conclusion

When I was a child, I went to an amusement park that had one of those mirrored mazes, and it included some mirrors that were able to cause you to see things upside down. I remember how I stumbled through the maze, quite disoriented.

A buddy of mine went into it over and over, spending all of his allowance to go repeatedly into the mirror maze. He eventually could not only walk through the maze without any difficulty, he could run throughout the maze and not collide or trip at all. His repeated “training” allowed him to eventually master the reorientation dissonance.

It seems that we need to try and make sure that today’s Machine Learning and Deep Learning gets beyond the existing ultra-brittleness, especially regarding the poses or orientation of objects. For most people, they would be dumbfounded to find out that the AI system can be readily fooled or confused by merely reorienting or tilting an object.

Those of us in AI know that the so-called “object recognition” that today’s ANN and DNN are doing is not anything close to what humans are able to do in terms of object recognition.

Contemporary automated systems are still rudimentary. This could be an impediment to the advent of AI self-driving cars. Would we want AI self-driving cars to be on our roadways and yet their AI can become intentionally or unintentionally muddled about a driving situation due to the orientation of nearby objects? I think that’s not going to fly. The objects orientation poses problem is real and needs to be dealt with for real-world applications.

Copyright 2019 Dr. Lance Eliot

This content is originally posted on AI Trends.

 

Google AI Ethics Council In Disarray After First Week

Google recently appointed an external ethics council to deal with tricky issues in artificial intelligence. The group is meant to help the company appease critics while still pursuing lucrative cloud computing deals.

In less than a week, the council is already falling apart, a development that may jeopardize Google’s chance of winning more military cloud-computing contracts.

On Saturday, Alessandro Acquisti, a behavioral economist and privacy researcher, said he won’t be serving on the council. “While I’m devoted to research grappling with key ethical issues of fairness, rights and inclusion in AI, I don’t believe this is the right forum for me to engage in this important work,’’ Acquisti said on Twitter. He didn’t respond to a request for comment.

On Monday, a group of employees started a petition asking the company to remove another member: Kay Coles James, president of a conservative think tank who has fought against equal-rights laws for gay and transgender people. More than 500 staff signed the petition anonymously by late Monday morning local time.

Employee activism on equal pay for women, sexual harassment, AI ethics and doing business in China has roiled Google over the past year. The protests have been effective. The company pulled out of a deal with the U.S. military to build object recognition technology for drones, prompting criticism from politicians including U.S. President Donald Trump. Google changed its policies after thousands of workers walked off their jobs to protest how the company deals with sexual harassment complaints.

At the same time, conservative politicians have said the company’s algorithms and content moderators unfairly discriminate against them in search results and YouTube, though there is no proof for these claims. Appointing Coles James, who is the president of influential right-wing think tank the Heritage Foundation, could be seen to conservatives as a sign that Google is hearing their concerns.

Some AI experts and activists have also called on Google to remove from the board Dyan Gibbens, the CEO of Trumbull Unmanned, a drone technology company. Gibbens and her co-founders at Trumbull previously worked on U.S. military drones. Using AI for military uses is a major point of contention for some Google employees.

Read the source article at Bloomberg

Warner Music Group Makes Record Deal with an Algorithm

One of the newest additions to the group of artists working with Warner Music Group — taking a spot alongside names like Ed Sheeran, Madonna, Coldplay and Camilla Cabello — is a bundle of code. It’s under contract to release 20 albums this year.

The creator of the algorithm is sound startup Endel, which uses artificial intelligence to make personalized audio tracks aimed at boosting people’s mood or productivity. Endel launched a little over a year ago in Europe, then became one of the up-and-coming companies chosen for Techstars Music’s startup accelerator program in 2018. Two months ago, Warner’s newly created Arts Music division agreed to tap Endel’s technology to make the world’s first-ever record label contract with an algorithm, covering distribution and publishing. The record company has released five of the 20 albums so far, as a collection of “sleep soundscapes” (Clear Night, Rainy Night, Cloudy Afternoon, Cloudy Night and Foggy Morning) designed to reduce anxiety; they’re all available for listening on services like Apple Music and Spotify.

As a company, Endel — which counts Amazon’s Alexa Fund, Japanese conglomerate Avex Inc and Major Lazer’s Jillionaire among its investors — actually focuses on creating tailor-made “custom sound frequencies” based on personal user inputs such as time of day and location, as well as biometric details such as heart rate. “We want to understand the context of your day and rearrange the whole environment around you,” Oleg Stavitsky, co-founder and CEO of Endel, tells Rolling Stone. Endel’s core algorithm takes thousands of sounds and assembles them into different templates based on the user inputs. The company launched a feature within Amazon’s Alexa Skills Store this week that allows users to receive custom sounds through Alexa-enabled devices, and its ultimate ambitions are even bigger: Stavitsky says he envisions an interconnected hardware-software ecosystem through which Endel can keep tabs on the rhythm of users’ daily lives via metrics like their driving patterns and the number of events on their calendar, then automatically creates a custom soundscape at the end of the day that helps them best unwind.

The record deal is just a nice side gig. “Warner approached us and we were hesitant at first because it counters what we’re doing here,” Endel’s co-founder and sound designer Dmitry Evgrafov tells Rolling Stone. “Our whole idea is making soundscapes that are real-time and adaptive. But they were like, ‘Yeah, but can you still make albums?’ So we did it as an experiment. When a label like Warner approaches you, you have to say ‘Why not.’” The other 15 records on the contract are themed around focus, relaxation and “on-the-go” modes and will roll out over the course of the year. All 20 albums will come out of Endel’s core algorithm, so they were technically, as Evgrafov says, “all made just by pressing one button.”

In a press release earlier this year, Warner’s Arts Music president Kevin Gore said he first learned of Endel through Techstars Music, and was immediately interested in having the team work on something for Warner. “Their innovative compositions provide unique listening experiences that will be introduced to a larger audience through the extensive reach of the Arts Music division’s marketing and distribution resources,” Gore said.

Read the source article in Rolling Stone.

Product Roundup: DDN, AI-powered Drug Discovery; Frictionless Facial Authentication

By AI Trends Staff

This is the first of a regular AI Trends round up of new products, partnerships and announcements in artificial intelligence for business. This month, we highlight new offerings and tools from DDN, dotData, Daedalean, and Alcatraz AI.

dotData, the first and only company focused on delivering end-to-end data science automation and operationalization for the enterprise, today announced the availability of Version 1.4 of its dotData Data Science Automation Platform. This latest update adds significant enhancements to the platform and provides users with deeper insights, increased flexibility, ease-of-use, and greater performance to meet their specific business goals. Key updates of the dotData Platform Version 1.4 include feature engineering from geo-temporal data, new State-of-the-Art Machine Learning Algorithms, Enhanced Automatic Data Preprocessing, and Drag-and-Drop Data Collection. Press Release

Searching for a way to ensure safety in various weather and emergency conditions, the founders of the Swiss company Daedalean decided to develop a universal software solution for an advanced autopilot to be used in eVTOLs (electronic vertical takeoff and landing vehicles). The software solution is powered by UNIGINE 3D engine for advanced autopilots used in eVTOL. Using virtual environment for computer vision training before trying a vehicle on real roads has a list of benefits. You can create various emergency scenarios without risk and high expenses, and safely imitate situations of AI-Human interaction. You can easily reconfigure a virtual training environment as many times as you need without monetary and time costs for the real polygon construction, and you can test your aircraft at a very early stage of engineering, without a ready-to-go prototype. Press Release

DataDirect Networks (DDN) has announced a partnership with Robovision, maker of a self-service deep learning platform, which enables organizations to bring concrete, maintainable AI applications live in mere weeks. The company also announced a new Platinum Tier within the DDN PartnerLink program, exclusively for NVIDIA DGX-certified resellers of DDN A³I with NVIDIA DGX-1, DGX POD, and DGX-2. This specialized tier is designed to provide added support to the resellers who work closely with AI and deep learning customers and best understand their needs, and to support them with the highest quality of service. The new tier will provide resellers with exclusive offerings from DDN, including access to proof of concept units for qualified opportunities, special pricing for demo units, guaranteed margins for registered deals, and special pricing on DDN A³I with NVIDIA DGX systems. Press Release

Ono Pharmaceutical, a Japanese pharmaceutical research and development company, and twoXAR, an artificial intelligence (AI)-driven biopharmaceutical company, have signed a drug discovery research collaboration to jointly discover and develop novel, efficacious treatments to address unmet medical needs in a specific neurological disease. Under the agreement, twoXAR will use its proprietary AI technology to identify a set of lead compounds which demonstrate a novel mechanism of action and will be further optimized by Ono for potential drug candidates. twoXAR will also predict a set of hypotheses which suggest the efficacy and safety of such lead compounds for the therapy. twoXAR and Ono will select several compounds with their hypotheses from this set to test in further validation studies. Ono will retain exclusive rights to develop and commercialize the compounds obtained through this collaboration throughout the world, and in return, twoXAR will receive research and license fees from Ono as well as development and sales milestones. Press Release

Agenus has launched an AI based, big data analysis platform called Adaptive Learning Platform Systems or “ALPS” with cancer immunotherapy research.  ALPS is designed to provide new insights into cancer biology, predict responses to treatment, identify new targets for drug development, and help develop precision combination strategies for cancer patients. With the integration of Big Data and immunotherapy research, Agenus scientists can now identify previously unknown trends to determine “immune fitness” in patients. Newsletter

Alcatraz AI, a company working on facial authentication frictionless access control for the enterprise, announced its first AI-enabled product today: a full stack platform with custom hardware on the edge, software on-premises and in the cloud. The Palo Alto-based startup leverages facial recognition, 3D sensing, and machine learning to enable highly secure and multi-person frictionless access control. Alcatraz has raised close to $6M in total funding. Alcatraz’s first product has the first-in-the-industry instant one-factor authentication for multi-person in-the-flow sensing. It uses multiple built-in cameras for real-time 3D facial mapping and NVIDIA GPU-powered deep neural networks. This allows for multi-person facial authentication as well as the automatic enrollment of individuals using current access control methods, such as badging. Alcatraz can deploy at locations where thousands of employees require access to buildings, without burdening HR or IT with the task of enrolling employees into a new security system. Additional features of the Alcatraz platform include tailgate detection, real-time notifications with video and access control analytics. The platform enables efficient and guardless entry flow and is scalable, fast to deploy, and easily integrates with existing access control systems. Website

Google to pull plug on AI ethics council

The council, launched on March 26, was meant to provide recommendations for Google and other companies and researchers working in areas such as facial recognition software, a form of automation that has prompted concerns about racial bias and other limitations.

Microsoft joins tech race to clean up shipping with big data

Maritime ships, which transport around 90% of the world’s goods across the seas, generate about 3 percent of global carbon emissions.

Government tightens check against Chinese e-comm players escaping taxes

The government is now asking the post office and courier companies to monitor shipments from China.

Managing the Complete Machine Learning Lifecycle: On-Demand Webinar now available!

On March 7th, our team hosted a live webinar—Managing the Complete Machine Learning Lifecycle—with Andy Konwinski, Co-Founder and VP of Product at Databricks.

In this webinar, we walked you through how MLflow, an open source framework for the complete Machine Learning lifecycle, helps solve for challenges around experiment tracking, reproducible projects and model deployment. Specifically, we covered:

  • Open source MLflow and overview of its core components: MLflow Tracking, MLflow Project, and MLflow Model.
  • Managed MLflow and how it integrates with the Databricks Unified Analytics Platform (now available to all customers at no additional cost).

In particular, we showed you how to:

  • Keep track of experiments runs and results across ML frameworks.
  • Execute projects remotely on to a Databricks cluster, and quickly reproduce your runs.
  • Quickly productionize models using Databricks production jobs, Docker containers, Azure ML, or Amazon SageMaker.

We demonstrated these concepts using notebooks and tutorials from our public documentation so that you can practice at your own pace. If you’d like free access Databricks Unified Analytics Platform and try our notebooks on it, you can access a free trial here.

We received so many questions from you at the end of the webinar, that we decided to hold a live session on Apr 11, 11am PST – Understanding MLflow: Ask the Experts – with Co-founders Matei Zaharia, Chief Technologist, and Andy Konwinski, VP of Product to answer them live! 

Save your spot now to understand how MLflow works under the hood, get exclusive insights into our roadmap, and ask your questions!

 

--

Try Databricks for free. Get started today.

The post Managing the Complete Machine Learning Lifecycle: On-Demand Webinar now available! appeared first on Databricks.

Thursday, 4 April 2019

Announcing Databricks Runtime 5.3

We are proud to announce the release of Databricks Runtime 5.3, which includes several new features and improvements, including:

  • GA of Delta Time Travel
  • Public Preview of MySQL table replication to Delta
  • Optimized DBFS FUSE folder for deep learning workloads (Azure-only)

Databricks Delta Time Travel (Generally Available)

Delta Time Travel has graduated to general availability. It adds the ability to query a snapshot of a table using a timestamp string or a version, using SQL syntax as well as DataFrameReader options for timestamp expressions.

SELECT count(*) FROM events TIMESTAMP AS OF timestamp_expression
SELECT count(*) FROM events VERSION AS OF version

Time Travel has many use cases, including:

  • Re-creating analyses, reports, or outputs (for example, the output of a machine learning model), which is useful for debugging or auditing, especially in regulated industries.
  • Writing complex temporal queries.
  • Fixing mistakes in your data.
  • Providing snapshot isolation for a set of queries for fast changing tables.

 

For more details, see Query an older snapshot of a table (time travel), and Merge Into (Databricks Delta).

MySQL table replication to Delta (Public Preview)

Databricks Runtime 5.3 lets you stream data from a MySQL table directly into Delta for downstream consumption in Spark analytics or data science workflows. Leveraging the same strategy that MySQL uses for replication to other instances, the binlog is used to identify updates that are then processed and streamed to Databricks as follows:

  • Reads change events from the database log.
  • Streams the events to Databricks.
  • Writes in the same order to a Delta table.
  • Maintains state in case of disconnects from the source.

For more details, see MySQL Table Replication to Databricks Delta.

Optimized DBFS FUSE folder for deep learning workloads (Azure only)

A new FUSE mount optimized for data loading, model checkpointing, and logging from each worker to a shared storage location, file:/dbfs/ml provides high-performance I/O for deep learning workloads.

For details, see Prepare Storage for Data Loading and Model Checkpointing.

Additional improvements

Apart from the above, Databricks Runtime 5.3 also includes:

  • Notebook-scoped library improvements
  • New Databricks Advisor hints
  • Delta performance improvements
  • And more …

To learn more about the release, please see the Databricks Runtime 5.3 Release Notes.

--

Try Databricks for free. Get started today.

The post Announcing Databricks Runtime 5.3 appeared first on Databricks.

7 Innovative Uses of Clustering Algorithms in the Real World

Clustering algorithms are a powerful technique for machine learning on unsupervised data. The most common algorithms in machine learning are hierarchical clustering and K-Means clustering. These two algorithms are incredibly powerful when applied to different machine learning problems.
 
Both k-means and hierarchical clustering have been applied to different scenarios to help gain new insights into the problem. Before diving into the innovative uses of clustering algorithms, I will first share an overview of the two algorithms. 

What is unsupervised learning?

Before we get started, let me first introduce the concept of unsupervised learning. Unsupervised learning is where you train a machine learning algorithm, but you don’t give it the answer to the problem.

1) K-means clustering algorithm

The K-Means clustering algorithm is an iterative process where you are trying to minimize the distance of the data point from the average data point in the cluster.

2) Hierarchical clustering

Hierarchical clustering algorithms seek to create a hierarchy of clustered data points.
The algorithm aims to minimize the number of clusters by merging those closest to one another using a distance measurement such as Euclidean distance for numeric clusters or Hamming distance for text.

Here are 7 examples of clustering algorithms in action.

1. Identifying Fake News

Fake news is not a new phenomenon, but ...


Read More on Datafloq

Databricks Runtime 5.3 ML Now Generally Available

We are excited to announce the general availability (GA) of Databricks Runtime for Machine Learning, as part of the release of Databricks Runtime 5.3 ML. Built on top of Databricks Runtime, Databricks Runtime ML is the optimized runtime for developing ML/DL applications in Databricks. It offers native integration with popular ML/DL frameworks, such as scikit-learn, XGBoost, TensorFlow, PyTorch, Keras, Horovod, etc. In addition to pre-configuring these popular frameworks, DBR ML makes these frameworks easier to use, more reliable, and more performant.

Since we introduced Databricks Runtime for Machine Learning in preview in June 2018, we’ve witnessed exponential adoption in terms of both total workloads and the number of users. Close to 1000 organizations have tried Databricks Runtime ML preview versions over the past ten months. To meet the rapidly growing demand, we continued to improve our integration and testing to arrive at a robust cadence of updating and adding libraries in Databricks Runtime ML. In addition, Databricks offers optimized features to improve the experience using these frameworks for developers. The positive feedback from our customers has led us to make Databricks Runtime for Machine Learning generally available (GA).

Databricks Runtime ML is now generally available across all Databricks product offerings:

  • Azure Databricks
  • AWS cloud
  • GPU clusters
  • CPU clusters

To get started, you simply select the Databricks Runtime 5.3 ML from the drop-down list when you create a new cluster in Databricks:

Creating a ready to use ML environment with Databricks Runtime for ML

Advantages of Databricks Runtime for Machine Learning

Databricks Runtime for Machine Learning focuses on three key areas: usability, reliability, and performance.

Ease of Use

All of the libraries identified in “Tiered Libraries in Databricks Runtime ML” (see below) come pre-configured in Databricks Runtime ML. You can start developing machine learning applications right away without the need to configure the environments themselves.

In the Databricks Runtime 5.0 ML release, we introduced HorovodRunner, which makes it easy to use the distributed deep learning framework Horovod. One key challenge with Horovod is usability, as it requires you to share code and libraries across nodes, configure SSH, execute complicated MPI commands, and so on. HorovodRunner abstracts all of these complications by providing a simple API to allow you to easily leverage the benefits of Horovod. With a few lines of code change, you can easily migrate your single node deep learning training code to run in a Databricks cluster.

Many ML libraries are developed for single-node use cases. We are constantly evaluating popular ML libraries and looking for ways to make them easier to run in a distributed system. For example, we are currently working on a distributed hyperparameter tuning feature. Stay tuned.

Reliability

Machine learning is a rapidly evolving space, and we want to make the latest and greatest tools available in Databricks Runtime for Machine Learning. Each of the pre-configured ML libraries in Databricks Runtime ML regularly releases new versions. To stabilize our environment, Databricks engineering runs daily integration tests against Databricks Runtime ML and stress-tests all new libraries before integrating or updating existing libraries.

Since the release of Databricks Runtime ML Beta ten months ago, we continued to expand our test suites and incorporated feedback from almost 1000 organizations to make ML workflows run smoothly.

We’ve taken a focused approach to maintaining and updating libraries in Databricks Runtime ML. Based on customer demand and market trends, we identified a list of libraries as “top-tier” libraries. For these “top-tier” libraries, Databricks plans to make faster updates and provides advanced support. See details in the “Tiered Libraries in Databricks Runtime ML” section below. With robust testing & integration in place, we feel confident to regularly add and update popular ML libraries in future Databricks Runtime ML releases.

Finally, a major initiative for Databricks Runtime ML in 2019 is to allow you to customize your ML environment. We are working on solutions that would let you to easily cherry-pick just the right set of ML libraries to include in Databricks Runtime ML. A lighter environment could lead to additional improvement in stability. Please stay tuned.

Performance

Databricks Runtime ML includes performance improvements beyond what is available in “off the shelf” open-source versions of several libraries. In the Databricks Runtime 5.0 ML release, we made improvements to both Apache Spark MLlib logistic regression and tree classifiers. When running in Databricks Runtime for ML, we observed ~40% speed-up in Spark Performance Tests compared to Apache Spark 2.4.0.

The GraphFrames library in Databricks Runtime for ML also contains an optimized implementation. Starting in 5.0 ML, GraphFrames in Databricks Runtime ML runs 2-4 times faster and supports even bigger graphs compared to open-source GraphFrames. Graph queries will utilize Spark cost-based optimization (CBO) to determine the join orders if the underlying node and edge tables contain column statistics. This can lead to as much as 100 times speed up, depending on the workloads and data skew.

The improved performance is only available in Databricks. You can take advantage of the improved performance in both Databricks Runtime and Databricks Runtime for ML.

We’ve also improved cluster launch time by 25% by reducing image size.

Tiered Libraries in Databricks Runtime ML

The Databricks Runtime for Machine Learning includes a variety of popular ML libraries. The libraries are updated regularly to include new features and fixes. A subset of popular libraries are marked as top-tier libraries. For these libraries, Databricks provides a faster update cadence, updating to the latest upstream package releases with each runtime release (barring dependency conflicts). Databricks also provides advanced support, testing, and embedded optimizations for top-tier libraries. Databricks Runtime 5.3 for Machine Learning includes the following libraries:

Top-tier libraries:

  • TensorFlow / TensorBoard / tf.keras
  • spark-tensorflow-connector
  • PyTorch
  • Horovod / HorovodRunner
  • GraphFrames

Other provided libraries:

  • Keras
  • spark-xgboost
  • MLeap
  • scikit-learn
  • pandas
  • Deep Learning Pipelines for Apache Spark
  • TensorFrames

Blobfuse (Azure-storage-fuse) as Default in Azure

Databricks Runtime has a basic FUSE client for DBFS, a local distributed file system installed on Databricks clusters. This feature has been very popular as it allows local access to remote storage. However, the current implementation does not allow fast enough data access required for developing distributed applications.

In Databricks Runtime 5.3 ML, azure-storage-fuse (Blobfuse) is now the default in Azure Databricks. You can now leverage the high-performance azure-storage-fuse without applying init scripts.

Two additional features are in the plan to further improve data I/O in Databricks. We are working on making Goofys the default for AWS users, achieving feature parity across Azure and AWS platforms. In addition, we are working to improve DBFS FUSE for faster sequential reads/writes, handling large files, etc.

Private Preview: MLlib-MLflow Integration

Databricks Runtime 5.3 ML supports automatic logging of MLflow runs for models fit using PySpark MLlib tuning algorithms CrossValidator and TrainValidationSplit. Before 5.3 ML, if you wanted to track PySpark MLlib cross validation or tuning in MLflow, you would have to make explicit MLflow API calls in Databricks notebooks. With MLflow-MLlib integration, when you tune hyperparameters by running CrossValidator or TrainValidationSplit, parameters and evaluation metrics will be automatically logged to MLflow. You can then review how the tuning affects evaluation metrics in MLflow.

This feature is in private preview. Contact your Databricks sales representative to learn about enabling it.

Other Library Updates

We updated the following libraries in Databricks Runtime 5.3 ML:

  • Horovod 0.16.0
  • TensorBoardX 1.6
  • PyArrow 0.12.1 (including support for BinaryType data)
  • The Databricks ML Model Export API has been deprecated. Databricks recommends using MLeap instead, which provides broader coverage of MLlib model types

Read More

  • Read more about Databricks Runtime 5.3 ML for Azure Databricks and AWS.
  • Try the example notebooks for distributed deep learning training for Azure Databricks and AWS on Databricks Runtime 5.3 ML.

--

Try Databricks for free. Get started today.

The post Databricks Runtime 5.3 ML Now Generally Available appeared first on Databricks.

Wednesday, 3 April 2019

Cybersecurity for Artificial Intelligence Solutions: A Framework for keeping AI from misbehaving

Scheduled for May 9, 2019, 1 pm to 2:00 pm EDT

Cybersecurity for Artificial Intelligence Solutions: A Framework for keeping AI from misbehaving
Sponsored by: OODA, LLC

The security of Artificial Intelligence is an emerging discipline. This webinar will provide deep insights based on years of direct experience in cybersecurity and in the fielding of analytical solutions.

Learning Objectives

  • Participants will gain a deeper understanding of the technical and non-technical dimensions of practical AI as well as an understanding of how AI has been going wrong in operational systems.

Register Now.

Speaker:

Bob Gourley, Founder and CTO, OODA LLC
Bob Gourley is the co-founder and Chief Technology Officer (CTO) of the Cybersecurity and Artificial Intelligence consultancy OODA LLC. Bob previously founded Crucial Point LLC, a technology research and advisory firm. He is the publisher of CTOvision.com Bob collaborates in the operation of OODALoop.com.

Register Now.

Blockchain Requires Industry Collaboration: The Launch of INATBA

When the web was developed over 25 years ago, the technologies in place significantly lowered the cost of building a global company. Thanks to the internet, it has become possible to reach a large part of the global population simply from behind your computer. Those companies who first understood the power of the web, and managed to execute their vision correctly, are now the leading global monopolies we are so familiar with: we use Google for finding information, Facebook or WeChat for social activities, Amazon to shop and Apple for our hardware, etc.

But times are changing since Satoshi Nakamoto distributed a paper among a small group of cryptography enthusiasts. Fast forward 11 years, and the underlying technology of the proposed bitcoin is rapidly changing how we run our organisations. Blockchain is a fundamental technology that changes how we perform transactions, how we collaborate and how we build our organisations.

Knowing what blockchain is and how it can contribute to improving your organisation and supply chain is one thing; knowing how to develop a blockchain strategy is another thing altogether. Especially, because when data becomes immutable, verifiable and traceable, it affects other important concepts such as privacy, security and ownership. Therefore, industry ...


Read More on Datafloq

Security spring cleaning

Today we published a security advisory that mostly informs about issues in Jenkins plugins that have no fixes. What’s going on?

The Jenkins security team triages incoming reports both to Jira and our non-public mailing list. Once we’ve determined it is a plugin not maintained by any Jenkins security team members, we try to inform the plugin maintainer about the issue, offering our help in developing, reviewing, and publishing any fixes. Sometimes the affected plugin is unmaintained, or maintainers don’t respond to these notifications or the followup emails we send.

In such cases, we publish security advisories informing users about these issues, even if there’s no new release with a fix. Doing so allows administrators to make an informed decision about the continued use of plugins with unresolved security vulnerabilities. Today’s advisory is overwhelmingly such an advisory.

See a plugin you love on this list and want to help out? Learn about adopting plugins.

Ola to launch in London as UK rollout builds

Ola raised $300 million from Hyundai Motor Group in March, giving the ride-hailing startup a valuation of about $6 billion.

5 Ways How Artificial Intelligence Is Impacting the Automotive Industry

 

A few decades ago, AI was the stuff of science fiction novels, movies, and television shows. A fantasy. A figment of an overactive imagination. Isaac Asimov. Star Trek. The Terminator.

But now?

AI is already here or has been for quite some time. It's in your smartphone, intelligent home system, and smart car. Entire cities are embracing Artificial intelligence to handle road safety.

It's everywhere, and the technology behind it is game-changing. One of the industries that took full advantage of using AI is the automotive industry. The automobiles we now drive are safer, more efficient and loaded with cutting edge tech.

Let's take a look at how artificial intelligence is impacting the auto industry.

1. Auto Insurance

AI in auto insurance has been a significant boon for companies and drivers. By using AI in risk assessment, insurance companies won't have to rely on past incidents to set premiums. Everything is in big data, and AI will crunch the numbers to predict how safe a driver will be.

Insurance companies factor in everything, from your health to recent relationship woes. If you continue to drive like a lunatic, AI will catch it, and you won't get insured! All this collected data means that you can check every ...


Read More on Datafloq

The Seven Patterns of AI

Scheduled for April 30, 2019, 1 pm to 2:00 pm EDT

There are many ways in which AI is being applied – from chatbots, image recognition, autonomous vehicles, speech and text applications, and many more. Despite all these different applications, Cognilytica has observed that there are only seven different patterns in which AI is being implemented. AI Projects generally use one or more of these patterns to accomplish their objectives.

In this webinar, Cognilytica analysts Kathleen Walch and Ronald Schmelzer will explain the seven patterns, and how they relate to successful implementation of AI projects. The patterns explained are Hypersonalization, Autonomous Systems, Predictive Analytics & Decision Support, Conversational / Human Interaction, Patterns & Anomalies, Recognition, and Goal-Driven Systems. All AI projects or implementation utilize one or more of these patterns as part of the solution. Each pattern is implemented with its own machine learning / cognitive process. We’ll explain how to identify which patterns are being used in a project, and how to utilize emerging methodologies to implement those patterns with the greatest degree of success.

In this webinar we’ll go over what each pattern is, questions your organization needs to answer to make sure you’re correctly using the pattern, and provide example use cases for how these patterns are implemented at organizations. We’ll also explore emerging AI-centric methodologies such as the Cognilytica Cognitive Technology Process Management Methodology (CCPTMM) and how it relates to implementation of the seven patterns.

In this webinar you’ll learn:

  • The seven AI implementation patterns
  • The right questions you need to answer before starting an AI project
  • How to identity which pattern(s) to use to solve your particular problem
  • How methodologies fit into implementation of the seven patterns

Register now!

Kathleen Walch, BS International Business, Managing Partner and Principal Analyst, Cognilytica
Kathleen is a principal analyst, managing partner, and founder of Cognilytica, an AI research and advisory firm, and co-host of the popular AI Today podcast. She is a serial entrepreneur, savvy marketer, AI and Machine Learning expert, and tech industry connector. Kathleen spent many years as the Content and Innovation Director for TechBreakfast, the largest monthly morning tech meetup in the nation with over 50,000 members and 3000+ attendees at the monthly events across the US.  In addition, she is a SXSW Innovation Awards Judge and AI / Hardware Meetup organizer. As a master facilitator and connector, who is well connected in the technology industry, Kathleen regularly meets with innovators in key markets and gets the opportunity to see the latest and newest technologies from game changing companies.

Ron Schmelzer, BS Computer Science & Electrical Engineering MIT and MBA Johns Hopkins University, Managing Partner and Principal Analyst, Cognilytica
Ron is principal analyst, managing partner, and founder of the Artificial Intelligence-focused analyst and advisory firm Cognilytica, and is also the host of the AI Today podcast, SXSW Innovation Awards Judge, founder and operator of TechBreakfast demo format events, and an expert in AI, Machine Learning, Enterprise Architecture, venture capital, startup and entrepreneurial ecosystems, and more.  Prior to founding Cognilytica, Ron founded and ran ZapThink, an industry analyst firm focused on Service-Oriented Architecture (SOA), Cloud Computing, Web Services, XML, & Enterprise Architecture, which was acquired by Dovel Technologies in August 2011.

Register now!