Friday, 3 May 2019

Salvage Yard Privacy Tattletales by Scrapped AI Self-Driving Cars

By Lance Eliot, the AI Trends Insider

My latest rental car that I picked-up at the Chicago O’Hare airport was a treasure trove of information about prior renters. I could readily see data that had been imported into the car’s infotainment system that had come from at least six different smartphones. Lots of favored playlists of people that I didn’t know, but I now knew their taste in music.

Even scarier for those prior renters was that many of their contacts had also gotten transferred into the on-board systems of the car. Plus, via the built-in GPS tracking, I could see the specific locations and dates/times of where many of these prior travelers had gone while using the rental car. If I had been a nefarious person, it would have been possible to use all of this info in rather untoward ways. In my case, I was just curious to see what others had opted to leave behind.

When you get out of a rental car and drop it off at the airport or other destination, sometimes people leave behind quite a curious set of physical odds and ends.

The lost-and-found at a car rental office will typically have tons of sunglasses, which are a popular leave-behind, and likewise kids’ toys are another common leftover. In my own forgetfulness, I had one time left a charger cord and figured it was worth going back to the car rental agency to see if they had retrieved it from the car that I had rented.

The car rental clerk, upon listening to my pleading to find my charger cord, took me to their oversized bin of leftovers, which had a plethora of forgotten items, including surprisingly that there were numerous sets of keys. Keys upon keys, on key chains, on key rings, on carabiners, you name it. Obviously leaving behind your keys is a frequently forgotten item too (one wonders, don’t people miss those keys, and don’t they then surmise that they must have left their keys in that rental car they checked-in?).

Anyway, the quite helpful clerk let me rummage around in the bin (security lapse?). There were a lot of charger cords. I could not prove for sure which one was mine, since there were many that looked just like mine. In the end, I retrieved one that was the same as mine and went on my merry way. The lesson of seeing the full-to-the-brim bin was to always double-check and then triple-check my rental cars before I turn them in.

Digital Artifacts Leftover in Cars

In a modern world, we leave behind not just physical artifacts but also digital artifacts.

It is easy to pair your smartphone to the infotainment and GPS systems of a rental car. Apparently, a lot of people don’t add together two-plus-two and realize that what goes in won’t necessarily come out. When you pair your smartphone, depending upon the nature of the settings, you might be allowing the car’s on-board systems to slurp-up all kinds of info out of your handy cellphone.

I suppose that there are some people that are unaware of the transfer of data that occurs. They live in their own non-digital world or are just part of the unwashed of the digital realm. For some of them, the smartphone and Bluetooth are already somewhat magical, I suppose.

There are likely other people that know that it happens, yet perhaps assume that it will miraculously be erased for them. Perhaps it’s like the Mission Impossible movies and the inputted data will sizzle and disappear after it is done with your travel journey. This transferred data will self-destruct in ten seconds, good luck, Jim.

Certainly, the rental car agency could include as part of their rental car clean-up checklist the step of them resetting and blanking out the on-board systems that might have collected your private info. This added step would be nice for the renting public. You would never need to worry about it again, assuming that the rental agency actually did the erasure properly, and consistently, and without making any errors or omissions.

Admittedly, it would be an extra step for the rental firm, and if you are cost conscious as a rental agency, you might say that the cost of the labor to do this reset operation is going to be significant. I know it might seem trivial as an action and not seemingly labor intensive if the reset is setup as a one-click operation, but when you multiply doing this for the thousand and thousands of rental cars in a fleet, the labor becomes mind boggling. If it also added “wasted” time to turning around a rental car, this is another downside factor and means that your fleet of cars cannot be as efficiently put back onto the road to earn more rents.

FCC Provides a Warning About Digital Artifacts Leftovers

The Federal Communications Commission (FCC) has tried to forewarn people about the Bluetooth pairing dangers, including saying this: “If you connect your mobile phone to a rental car, the phone’s data may get shared with the car.  Be sure to unpair your phone from the car and clear any personal data from the car before you return it. Take the same steps when selling a car that has Bluetooth” (this is stated at the FCC’s web site, https://www.fcc.gov/consumers/guides/how-protect-yourself-online).

Note that the FCC warning also mentions the notion of erasing your data when you sell your own personal car.

I’m betting that many people neglect to do so. I know this for a fact because I bought a “previously owned” (let’s say it more plainly, “used”) car, and it had all sorts of data from the previous owner. What makes this more insidious is that the data covered a multi-year period of time. In the case of a rental car, presumably it would likely have less data, though it also depends upon how much has been retrieved from your smartphone.

I’m sure there are a lot of people though that when selling a car are more likely to think about their smartphone data that’s on-board the car. I would guess that it is something that you would be inclined to consider. A rental car is maybe less likely to be on your mind in terms of having paired your smartphone to it. For a car that you owned, you would be much more cognizant about having used the car’s on-board systems, it would seem.

Here’s an added twist for you, what about when your car gets wrecked and it is hauled off to a salvage yard.

Would you be thinking about the data that you’ve left in your wrecked car?

Probably not.

What Happens to a Wrecked Car

If you are car has gotten so wrecked that it cannot be repaired, the odds are that you are mentally and physically done with that car. The car is probably disfigured. It looks horrible. Nobody wants to try and deal with their now contorted and bruised car. It is like losing a loved one, in a sense, since we often grow fond of our cars. I am not saying it’s the same as a person or a pet, and merely trying to point out that many times we get emotionally attached to our cars.

I did so. One of my first cars was a nifty sports car. I had wanted that car for many years and saved up to buy it. I was pretty happy the day I bought it. Several years later, the car got stolen. I was devastated at first. My prized car was gone. I got angry. How dare thieves steal my car! I wanted revenge. I hoped the police would find the car thieves and, well, let’s just say that street justice seemed to be a fine way to deal with them, if you know what I mean.

I was told by the police that the odds were pretty high that the car was stolen by a local gang. The gang would joyride the car until they had enough of the fun or until the car itself was no longer able to run. Apparently, there was a rash of gang initiation rights that involved stealing a car, and my kind of sports car fit the profile of what was needed to get into a gang. Who knew?

After a few days of being despondent and hoping to get my car back, I gradually changed my mind about the matter. I didn’t want the car anymore. It had become soiled by the intrusion of the gang, if indeed that’s what had occurred. The car would never be the same, even if the gang somehow decided to park it someplace and walk away from it. My emotional attachment to the car became detached.

Amazingly, about ten days after the car had been stolen, I got a call from the police department, my car had been discovered. I went right away to go see it. The police had it at the official police impound. When my eyes saw the wreck that was left of my car, I knew at that moment that I probably should not have come to see it. It looked lifeless. Plus, the gang had driven it until a tire blew, and they kept driving on the rim, and ultimately rammed it into another car. The poor thing was a mangled and nearly unrecognizable variant of my prized sports car.

The gang had apparently zestfully stripped everything out of the interior once they had decided to abandon it. I mean everything was gone. There wasn’t much of anything inside leftover, not even the flooring mats and carpets. This sucker was picked clean. Imagine a skeleton of a car, prior to the auto maker putting the guts into it.

Believe it or not, I had never paired my smartphone to the car. I know this seems nuts, but it was just one of those get-around-to-it kinds of things that I had not done. Fortunately, there wasn’t any personal info in the car, other than the car registration had been in the glovebox, which was now gone, along with everything else that had been taken.

In discussing the car with the insurance company, they advised that the cost to repair the car would be excessive and recommended that the car be considered totaled. I quickly agreed. As mentioned, I was over the car by now. And, upon seeing it as a now picked over corpse, I could not imagine ever driving it again, in spite of whatever astounding repairs and fix-up that possibly could be undertaken.

That’s the last I ever saw of my sports car.

When I gently and hesitantly asked the insurance agent what would happen to the carcass (I’m wasn’t exactly sure that I wanted to know), he explained that it would be hauled off to a salvage yard. Some people call them junkyards, others refer to them as scrapyards. A rose by any other name. In the end, my car would be dismantled.

Any usable parts would be potentially resold onto the used parts market, or in some cases the scrapyard hangs onto the parts or to the carcass and allows prospective used-parts buyers to come and pick over the skeletons. There seems to be a thriving market of people needing to fix up cars and wanting to find the original parts that fit to the same brand and model of car that they own. Often these are car collectors.

My insurance agent explained that the cost of buying a new part from an auto maker is likely going to be a lot pricier than getting a used part that was once on a no-longer operating car. Some scrapyards remove the reusable parts and place them into a salvage warehouse, nearly arranged. More often, the scrapyards just pile up the “deceased” cars and allow car part seekers to roam around and find whatever they think they need.

He also explained that the unusable elements could be turned into scrap that can be sold at bulk prices, especially scrap metal. The front windshield was smashed, but the other windows were still intact, and so those could be removed and potentially sold as-is. My front bumper was ripped off the car entirely and the headlights were pretty much goners too. Meanwhile, the taillights seemed to still be workable, along with the mirrors, the exhaust system, and so on.

I decided that perhaps my sports car would make a better life for someone else, doing so by my “donating” them to the salvage yard. Well, okay, I didn’t actually donate the car, I instead got a check from the insurance company that covered the insured value. I just say in my own mind that I donated it, similar to providing donated organs for science. I’d like to imagine that my banged up, destroyed, tainted sports car had become a helpful source of parts that would make others happy.

The insurance agent told me that likely 75% of the average wrecked car can be put to some other use. I really had no idea what would happen to a wrecked car and the idea that it is “recycled” in this manner seemed generally impressive. Better than it all just sitting in a big heap and rotting away for years upon years.

Have you ever had a car that was scrapped?

According to statistics by the federal government, there are about 15 million cars per year in the United States that end-up in a scrapyard. There are an estimated 250 million cars in the United States. Thus, as you can see, only about 6% of the cars in-hand seem to go to scrapyards each year. I apparently am one of the “lucky” few to have it happen to their car. At least I wasn’t in the car during a car accident that ultimately might have wrecked the car and gotten it to go to the wrecking year. Having a gang steal it was an “easier” way to have it end-up at a salvage yard.

The Case of AI Self-Driving Cars

What does this have to do with AI self-driving cars?

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. One aspect that few of the auto makers or tech firms are considering is what will happen to the data that’s on-board an AI self-driving car once the self-driving car ends-up in a salvage yard.

Allow me to elaborate.

I’d like to first clarify and introduce the notion that there are varying levels of AI self-driving cars. The topmost level is considered Level 5. A Level 5 self-driving car is one that is being driven by the AI and there is no human driver involved. For the design of Level 5 self-driving cars, the auto makers are even removing the gas pedal, brake pedal, and steering wheel, since those are contraptions used by human drivers. The Level 5 self-driving car is not being driven by a human and nor is there an expectation that a human driver will be present in the self-driving car. It’s all on the shoulders of the AI to drive the car.

For self-driving cars less than a Level 5, there must be a human driver present in the car. The human driver is currently considered the responsible party for the acts of the car. The AI and the human driver are co-sharing the driving task. In spite of this co-sharing, the human is supposed to remain fully immersed into the driving task and be ready at all times to perform the driving task. I’ve repeatedly warned about the dangers of this co-sharing arrangement and predicted it will produce many untoward results.

For my overall framework about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

For the levels of self-driving cars, see my article: https://aitrends.com/selfdrivingcars/richter-scale-levels-self-driving-cars/

For why AI Level 5 self-driving cars are like a moonshot, see my article: https://aitrends.com/selfdrivingcars/self-driving-car-mother-ai-projects-moonshot/

For the dangers of co-sharing the driving task, see my article: https://aitrends.com/selfdrivingcars/human-back-up-drivers-for-ai-self-driving-cars/

Let’s focus herein on the true Level 5 self-driving car. Much of the comments apply to the less than Level 5 self-driving cars too, but the fully autonomous AI self-driving car will receive the most attention in this discussion.

Here’s the usual steps involved in the AI driving task:

  •         Sensor data collection and interpretation
  •         Sensor fusion
  •         Virtual world model updating
  •         AI action planning
  •         Car controls command issuance

Another key aspect of AI self-driving cars is that they will be driving on our roadways in the midst of human driven cars too. There are some pundits of AI self-driving cars that continually refer to a utopian world in which there are only AI self-driving cars on the public roads. Currently there are about 250+ million conventional cars in the United States alone, and those cars are not going to magically disappear or become true Level 5 AI self-driving cars overnight.

Indeed, the use of human driven cars will last for many years, likely many decades, and the advent of AI self-driving cars will occur while there are still human driven cars on the roads. This is a crucial point since this means that the AI of self-driving cars needs to be able to contend with not just other AI self-driving cars, but also contend with human driven cars. It is easy to envision a simplistic and rather unrealistic world in which all AI self-driving cars are politely interacting with each other and being civil about roadway interactions. That’s not what is going to be happening for the foreseeable future. AI self-driving cars and human driven cars will need to be able to cope with each other.

For my article about the grand convergence that has led us to this moment in time, see: https://aitrends.com/selfdrivingcars/grand-convergence-explains-rise-self-driving-cars/

See my article about the ethical dilemmas facing AI self-driving cars: https://aitrends.com/selfdrivingcars/ethically-ambiguous-self-driving-cars/

For potential regulations about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/assessing-federal-regulations-self-driving-cars-house-bill-passed/

For my predictions about AI self-driving cars for the 2020s, 2030s, and 2040s, see my article: https://aitrends.com/selfdrivingcars/gen-z-and-the-fate-of-ai-self-driving-cars/

Wrecked AI Self-Driving Cars Are a Data Treasure Trove

Returning to the topic of AI self-driving cars that end-up in a salvage yard, let’s consider why this might happen and what makes it different from a conventional car that is hauled into such a resting place.

First, the big reason that an AI self-driving car differs from a conventional car in terms of the salvage yard is that an AI self-driving car is chock full of sensors and computer processors.

A conventional car is likely to have a limited set of sensors, often not nearly as powerful and full-bodied as those that would be used on a true AI self-driving car. And, the computer processors in a true self-driving car need to be top-of-the-line, superfast to handle the AI running aspects, more so than the processors on a conventional car.

I am not saying that today’s modern conventional cars don’t have some semblance of sensors and processors. Instead, I am pointing out that on a Level 4 or Level 5 self-driving car, the odds are they are a step-up in terms of capabilities, along with often higher costs too, at least when purchased new.

Furthermore, the amount of on-board system memory is likely a lot more than you would have on a conventional car.

This is where the concern really focuses about having your wrecked AI self-driving car towed into a salvage yard. Remember my earlier story about car rentals that are turned-in and the renter has left personal data in the on-board systems? Magnify that kind of leftover info a thousand-fold, and you have the situation we are facing with AI self-driving cars.

An AI self-driving car is likely to have captured video streams that are left intact in the wrecked AI self-driving car. There is a treasure trove of telematic data about the activity of the self-driving car. There could be data that was transmitted back-and-forth via the OTA (Over-The-Air) electronic communications that might have taken place between your self-driving car and the cloud of the auto maker. There could be V2V (vehicle-to-vehicle) electronic communications stored in the on-board systems, involving your self-driving car communicating with other self-driving cars.

All of this then is in addition to whatever you might have placed into the self-driving car via your connected smartphone.

Things get even worse.

If your AI self-driving car has a voice activated Natural Language Processing (NLP) system that allows you to give verbal commands to the self-driving car, those might also be stored in the on-board systems. If the self-driving car is a Level 2 or Level 3, in which you co-shared the driving task, the odds are that there might be captured info about your driving and the driving aspects of the AI system.

Tesla Examples Found by Researchers

Let’s consider the Tesla cars.

According to Tesla’s owner manual, here’s the kind of Telematics info that could be kept on-board the car:

“To improve our vehicles and services for you, we may collect certain telematics data regarding the performance, usage, operation, and condition of your Tesla vehicle, including: vehicle identification number; speed information; odometer readings; battery use management information; battery charging history; electrical system functions; software version information; infotainment system data; safety-related data and camera images (including information regarding the vehicle’s SRS systems, braking and acceleration, security, e-brake, and accidents); short video clips of accidents; information regarding the use and operation of Autopilot, Summon, and other features; and other data to assist in identifying issues and analyzing the performance of the vehicle.” (source: https://www.tesla.com/about/legal).

Plus, this kind of data too:

“Data about accidents involving your Tesla vehicle (e.g., air bag deployment and other recent sensor data); data about remote services (e.g., remote lock/unlock, start/stop charge, and honk-the-horn commands); a data report to confirm that your vehicle is online together with information about the current software version and certain telematics data; vehicle connectivity information; data about any issues that could materially impair operation of your vehicle; data about any safety-critical issues; and data about each software and firmware update.”

In case you are thinking that this is merely an abstract problem and would not occur in the real-world, there is a fascinating study that was recently released about a computer security company that bought some wrecked Tesla cars at a salvage yard and examined those cars to see what they could find (for an article and a video of what they found, see: https://www.cnbc.com/2019/03/29/tesla-model-3-keeps-data-like-crash-videos-location-phone-contacts.html).

The researchers pored into four cars that they obtained, specifically a Tesla Model X, a Tesla Model S, and two of the Tesla Model 3 cars. Of course, they found paired data from smartphones. I’d say that’s pretty much to be expected of any modern-day car, and not especially surprising or unusual. This included nearly a dozen phonebooks of contact info, and various GPS navigation locations.

What’s more interesting is the aspect that for one of the Model 3’s, the researchers extracted the video of the Model 3 of when it had crashed. The car had veered off the road and crashed, which the front cameras recorded. Tying this to the GPS data, the researchers could ascertain the location, Orleans, Massachusetts, occurring on Manequoit Road, and the day and time of the crash. The airbags also deployed. They also tied the crash to the smartphone that was plugged into the car at the time, being able to figure out presumably the person driving the car.

They also looked at the log of the phone use and could see that a phone call from a family member (a contact in the database) had called the driver of the car, moments before the crash occurred.

I think we would all be rather shocked to find out that our private details could be so easily gleaned from our wrecked car. You would normally likely assume that those kinds of details would need to be gotten by a court order or a subpoena of some kind.

Also, you would likely assume that the data would be secured in some manner, making it hard for just anyone to retrieve. According to the researchers, by-and-large the data collected was unencrypted. There was no need to try and crack any difficult ciphers or codes.

I don’t want to seemingly be picking on Tesla, and it should be pointed out that the Tesla licensing does have some warnings about a salvaged Tesla, including this:

“An unsupported or salvaged vehicle is a vehicle that has been declared a total loss, commonly after extensive damage caused by a crash, flooding, fire, or similar hazard, and has been (or qualifies to be) registered and/or titled by its owner as a salvaged vehicle or its equivalent pursuant to local jurisdiction or industry practice. Salvage registration/titling typically can never be removed from the vehicle so that all future persons understand the condition and value of the vehicle. Tesla does not warrant the safety or operability of salvaged vehicles. Repairs performed to bring a salvaged vehicle back into service may not meet Tesla standards or specifications and that is why the vehicle is unsupported. Consequently, any failures, damages, or injuries occurring as a result of such repairs or continued operation of an unsupported vehicle are solely the responsibility of the vehicle owner” (source: https://www.tesla.com/about/legal).

In a manner of speaking, presumably it is the duty of the car owner to cope with the matter of having their own car salvaged and taking any needed steps.

According to the researchers, Tesla apparently reported to them that:

“Tesla already offers options that customers can use to protect personal data stored on their car, including a factory reset option for deleting personal data and restoring customized settings to factory defaults, and a Valet Mode for hiding personal data (among other functions) when giving their keys to a valet. That said, we are committed to finding and improving upon the right balance between technical vehicle needs and the privacy of our customers” (as stated in: https://www.cnbc.com/2019/03/29/tesla-model-3-keeps-data-like-crash-videos-location-phone-contacts.html).always

You can interpret the response by Tesla as befits your own views about what responsibility the car maker has versus the car owner.

Coping With AI Self-Driving Cars Once Wrecked

As the advent of AI self-driving cars continues to increase, there will be more and more circumstances involving wrecked AI self-driving cars.

Right now, the Teslas, which are considered pretty much a Level 2, those are the most prevalent of any semblance of an AI self-driving car and so it is logical that those would be getting wrecked, in the normal course of being on the roads, and end-up in salvage yards.

With the emergence of Level 3’s, once those become relatively popular, they will ultimately get into wrecks, sorry to say, but it’s a fact, because they are cars, and that’s what happens with cars, and so those too will eventually get piled into scrapyards.

The Level 4 and Level 5 self-driving cars are right now working in experimental modes and prove-of-concept (POC) modes, and are not owned by individuals per se. Instead, they are being crafted by auto makers and tech firms. This means those self-driving cars are lovingly tended by a slew of expert mechanics and AI professionals. If those self-driving cars get into a wreck, it isn’t as though they will just tow the self-driving cars to the nearest salvage yard and junk them there.

Nope. Those babies lead a pampered life, right now.

My point being that the auto makers and tech firms have not had to deal with the end-of-life aspects as yet of self-driving cars. We are still so much at the start of the life-cycle that thinking about the end of the life cycle is nearly unimaginable. AI developers that I talk with are oft to scoff at the end-of-life of their creations, doing so because they are harried and knee deep into just trying to make AI self-driving cars that work, being able to have the AI drive around without hitting anything or anyone.

For my article about reverse engineering of AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/reverse-engineering-and-ai-self-driving-cars/

For what happens when an AI self-driving car gets into an accident, see: https://www.aitrends.com/selfdrivingcars/accidents-happen-self-driving-cars/

For the burned-out AI developer aspects, see my article: https://www.aitrends.com/selfdrivingcars/developer-burnout-and-ai-self-driving-cars/

For my article about the egocentric AI developers’ aspects, see: https://www.aitrends.com/selfdrivingcars/egocentric-design-and-ai-self-driving-cars/

You might be tempted to suggest that at least the on-board data should always be encrypted.

By doing so, it would mean that even if the wrecked self-driving car was given to a salvage yard, it would be arduous or perhaps infeasible for anyone to readily pluck the data out of the car in terms of knowing what the data actually contained (they might be able to grab it, but it would appear to be undecipherable).

Though this is a good idea, it also offers the downside of having to be continually encrypting and potentially decrypting data to make use of it to drive the self-driving car by the AI system.

This means that the on-board computer systems are going to do a lot of added computational work. The data being collected by the sensors would need to be turned from plaintext or plain-data into encrypted data. Would this happen only once the data is stored? That’s data in-rest or in-place. Would it also occur when the data is flowing throughout the on-board system, which is data-in-motion?

There are lots of questions to be considered. Would the added computational effort dilute the on-board computational processors and distract those processors from the “real work” of running the AI to drive the car? Would the time it takes to encrypt and decrypt create a potential delay in having the AI be able to readily make driving decisions, which are real-time and life-or-death kinds of matters?

Some say that maybe have the data encrypted at the end of a driving day, thus only the data that might so happen to be “live” when a wreck occurs would be potentially unencrypted.

For more about the cognition timing aspects, see my article: https://www.aitrends.com/selfdrivingcars/cognitive-timing-for-ai-self-driving-cars/

For the backdoor security matters, see: https://www.aitrends.com/selfdrivingcars/ai-deep-learning-backdoor-security-holes-self-driving-cars-detection-prevention/

For the role of Event Data Recorders, see my article: https://www.aitrends.com/selfdrivingcars/event-data-recorders-edr-self-driving-car-need-black-box-relook/

For my article about privacy concerns of AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/privacy-ai-self-driving-cars/

Another suggestion is that the self-driving car should have a “wrecked mode” that would automatically kick-in when the self-driving car gets into a crash of some kind. This would either encrypt the data at that juncture, though you need to hope that the processors and systems are working sufficiently that this could actually occur after the crash has happened, or the wrecked mode might erase everything, similar to my Mission Impossible comment earlier (again, this assumes that the AI is still working sufficiently).

One concern about the erasing of data would be whether the data might be needed for purposes of establishing any legal claims about a crash that has occurred. Whether or not our society would allow the auto makers or tech firms to summarily have a feature that would automatically erase everything, well, that’s a pretty big if.

You could say that it is up to the owner of the self-driving car to take proper action with their wrecked car. Thus, if someone is “stupid enough” to handover their wrecked car to a salvage yard, and leave all of their personal data in it, that’s their own act of being a dolt.

Some would have more sympathy toward the owner of a wrecked car. Would the owner understand that it is their responsibility to deal with the data? Would they realize that the data was even being collected? Would they realize that it wasn’t automatically being encrypted for them? Would they understand that it is something they need to take overt action about?

I think we can likely agree that having something in an owner’s manual is not quite the most broadcast way to inform car owners. How many of us actually read the owner’s manual? It is akin to those that download and use an app, which has a 50-page online licensing contract, and for which most people just click yes and agree to the terms. If the app then gives up all their personal data and sells it to the dark web, do we merely say that those people were dolts?

It could be that some might argue that the salvage yards have an obligation to not allow the data from the towed-in self-driving cars to be handed out. Perhaps their should be legislation that requires salvage yards to protect your data and inform you about it. I doubt that many salvage yards will welcome such an added burden onto their shoulders.

You might say that it should be on the shoulders of the auto maker and tech firms that make the AI self-driving cars. That’s again something that has yet to be ascertained in terms of what the range and nature of their duties are. Much of this is still an open market approach and there is little yet in the way of regulatory rules about it.

I would guess that we’ll likely see lawsuits that will also arise due to these matters. Someone that has had a wrecked AI self-driving car that reveals private aspects will launch a lawsuit against the auto maker or tech firm, perhaps at the insurance firm, perhaps at the salvage yard, and maybe at anyone or anything in the life cycle steps after a self-driving car has gotten wrecked.

For my article about responsibilities and AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/responsibility-and-ai-self-driving-cars/

For federal regulations about AI self-driving cars, see my article: https://www.aitrends.com/selfdrivingcars/assessing-federal-regulations-self-driving-cars-house-bill-passed/

For whether we are creating a kind of Frankenstein, see: https://www.aitrends.com/selfdrivingcars/frankenstein-and-ai-self-driving-cars/

For the soon to emerge lawsuits bonanza, see my article: https://www.aitrends.com/selfdrivingcars/first-salvo-class-action-lawsuits-defective-self-driving-cars/

Things Will Get Worse In Terms of What’s On-Board

I’ll add more fuel to the fire.

It seems likely that true Level 5 AI self-driving cars will have cameras pointing inward and be recording the audio and video of whatever happens inside of the self-driving car. Why this kind of intrusion? It can be to help the AI figure out what the human passengers are doing and what they want the AI to do for them.

There’s another equally practical reason, namely for ridesharing purposes. Most would agree that the AI self-driving car of a Level 5 will be used for ridesharing purposes. Even if you own your own Level 5 self-driving car, you will likely let it roam and be a ridesharing vehicle while you are at work or asleep, allowing your self-driving car to make money for you.

By having the cameras that point inward, you can keep track of those pesky ridesharing passengers that might decide to trash the inside of your shiny AI self-driving car. Or, perhaps it could be that someone is having a heart attack and needs urgent help, which the AI might be able to detect by scanning the interior video and then contacting 911 or routing the self-driving car to the nearest hospital.

The overall point is that this kind of private data would also be presumably kept on-board the self-driving car. Once again, it might be accessed once the self-driving car is relegated to a salvage yard, if not otherwise protected or erased.

I’ll scare you about the outward facing cameras too.

As your AI self-driving car goes down the street in your neighborhood, it is capturing video, along with possibility audio, and radar, and LIDAR, and ultrasonic waves, which could be kept on-board the self-driving car. It might be sitting in there, a view of all of your neighbors, their dogs and cats, their comings and goings.

When you park your AI self-driving car in your garage, it might still be recording. This could occur in that the AI self-driving car might be setup to wait for you to ask it to do something, so it is sitting there in a semi-alert fashion. It is akin to Alexa or Siri, listening for a prompting word. Though in theory the listening mode is not recording, you never know how it might really have been established.

Some believe that the AI of the self-driving car will be a kind of therapist, allowing you or your children to interact with it on your daily commute. The AI might try to help you with that problem at work, or difficulties with your spouse. Or, your children might confide that they are failing in their classes and want to run away from home. All of this potentially could be recorded by the AI system.

This AI would use a mixture of NLP, socio-behavioral techniques, and possibly Machine Learning and Deep Learning. Whatever methods or technologies used, it all depends upon having data, including collecting it and keeping it around, in some manner, whether in whole or in a compressed or selected manner.

We really haven’t as yet established what the boundaries are going to be about the recording of such data.

I know some pundits claim that the voluminous data is so voluminous that it would not make any sense for the self-driving car to keep it on-board. The amount of on-board computer memory would be overly costly, use up too much power, and be large and heavy, weighing down the AI self-driving car. They say that this data either won’t be kept, or it will be shunted up to the cloud via the OTA.

We’ll have to wait and see how this plays out.

For my article about Machine Learning and AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/machine-learning-ultra-brittleness-and-object-orientation-poses-the-case-of-ai-self-driving-cars/

For the use of socio-behavioral techniques and AI, see: https://www.aitrends.com/features/socio-behavioral-computing-for-ai-self-driving-cars/

For the rise of ridesharing and AI self-driving cars, see my article: https://www.aitrends.com/selfdrivingcars/ridesharing-services-and-ai-self-driving-cars-notably-uber-in-or-uber-out/

For my article about the affordability of AI self-driving cars, see: https://www.aitrends.com/selfdrivingcars/affordability-of-ai-self-driving-cars/

For my article about how IoT plays into this, see: https://www.aitrends.com/selfdrivingcars/internet-of-things-iot-and-ai-self-driving-cars/

Conclusion

If I was sad when my sports car went to the salvage yard, imagine how I might feel when my true Level 5 AI self-driving car (of the future) ends-up there too. My sports car could not interact with me, and yet I considered it my friend. For the Level 5 self-driving car, presumably it will be a friend, a confidant, a father confessor, a butler, and probably know more about me than any other living human being. Yikes!

In any case, we do need to all start considering what to do about AI self-driving cars and the data they are going to be collecting. The focus herein was what happens to the data when the self-driving car gets junked to a salvage yard.

That’s just the tip of the iceberg. While the self-driving car is still fully active, we need to be worrying about the data and how it is being collected and who can access it.

The next time you drive past a scrapyard, look at the pile of cars, and think to yourself about the hidden secrets that will someday be there, embedded into the computer memory of those AI self-driving cars that were unceremoniously dumped there. Perhaps we’ll prevent that data dumpster treasure trove from happening, if we take heed now in the design and development of AI self-driving cars.

Copyright 2019 Dr. Lance Eliot

This content is originally posted on AI Trends.

 

AI Adoption on the Upswing; Investments Increasing

Enterprise adoption and attitudes: Some progress, some FOMO

Some 25% of businesses surveyed have implemented cognitive technologies such as AI or machine learning, either as pilot projects or as long-term strategies; 41% are using Robotic Process Automation (RPA) extensively or across multiple functions, 26% of respondents are using robotics, 22% are using AI; 64% saw growth ahead in robotics, 80% predicted growth in cognitive technologies, and 81% predicted growth in AI (Deloitte 2019 Global Human Capital Trends).

87% of companies are adopting at least one transformative technology (video everywhere, the internet of things (IoT), artificial intelligence (AI), 5G, blockchain and the cloud) in their businesses; only 26% believe appropriate business models are in place to capture full value from these technologies (HIS Markit Digital Orbit).

More than 30% said their companies have allocated $50 million or more to smart automation projects, and more than half have already spent at least $10 million; the initiatives include various combinations of robotic process automation (RPA), artificial intelligence, machine learning, cognitive computing and analytics; highest expenditure levels were for the finance and accounting category, marked by 23% of respondents as receiving investment of slightly more than US$50 million; the technology that organizations are experimenting with or piloting the most is AI (36); 30% of companies are opting not to invest or are unsure of their plans for smart automation (KPMG Easing the Pressure Points).

85% of airlines are planning to use AI for virtual agents and chatbots by 2021; 79% of airports are currently using, or planning to use, AI for predictive analysis to improve their operational efficiencies. (SITA How strong investment in digitization will transform the passenger journey).

Lawyers surveyed think AI will be valuable for tracking billable time (53% of US layers, 49% of UK lawyers), conflicts clearance (43% and 41%), and compliance with client billing documents (34% for both US and UK lawyers) (Intapp survey reveals lawyers’ attitudes toward technology).

The portion of auto companies not using or testing AI rose to 39% in 2019 from 26% in 2017 (Capgemini).

“The accelerated growth of RPA is being driven by high levels of efficiency and productivity that can now be achieved from intelligent automation, which combines advanced RPA, artificial intelligence and embedded analytics. The demand for RPA solutions has surged as legacy companies are now competing with ‘digital native’ companies like Amazon and Uber, in which nearly every part of the business is completely automated”—Mihir Shukla, CEO of Automation Anywhere Inc., an RPA maker that expects to deploy three million software robots at organizations worldwide by 2020, a 200% increase from today (Wall Street Journal).

Consumer adoption and attitudes: Make us trust the machine

64% of US consumers will not buy self-driving cars, 63% will not spend more for self-driving features; two-thirds of survey respondents said self-driving cars should be held to higher government safety standards than traditional vehicles driven by humans (Reuters/Ipsos Americans still don’t trust self-driving cars).

Apple researchers tested people against three types of virtual assistant: a chatty system, a non-chatty system, and one which tried to mirror the chattiness of the user. The study found that people tend to prefer chatty assistants to non-chatty ones, and have a significant preference for agents whose chattiness mirrors the chattiness of the human user, as “interacting in this fashion increases feelings of trust” (Import AI #141 and Mirroring to Build Trust in Digital Assistants).

Read the source article in Forbes.

Real World Impact of AI on Labor, Economy Told Through Case Studies

By Peter Lo, Partnership on AI

The  impact of artificial intelligence (AI) on the economy, labor, and society has long been a topic of debate — particularly in the last decade — amongst policymakers, business leaders, and the broader public. Estimates of its current and imminent labor and productivity impacts have varied widely, often reaching contradictory conclusions.

A salient question for researchers, managers, and economic analysts has been whether large investments in AI and machine learning (ML) are warranted. Can the promises of AI be realized,  and if so, what are their potential impacts on the various stakeholders involved?

To  help elucidate these various areas of uncertainty, the Partnership on AI (PAI) is publishing a series of case studies conducted by the PAI’s Partnership on AI Working Group on “AI, Labor, and the Economy” (AILE).

The AILE  Case Study Compendium investigates the labor  implications and productivity impacts of AI implementation through a series of case studies across different applications, geographies, and sectors. Using interview-based methods, we examined the impact of AI applications at three companies: Axis Bank, Tata  Steel Europe (TSE), and Zymergen.

Three common themes emerged:

Successful adoption of new AI systems required buy-in from management and the workforce alike. In doing so, having intelligible ML models was of paramount importance.

At Zymergen, a young company with AI/ML  central to its founding mission, employee buy-in was generally assumed from the beginning. At TSE and Axis Bank, however, teams needed to approach changes in culture and practice more proactively with relevant internal stakeholders.

In all three cases, executive  support and early investment in employee training and awareness-building helped to engender trust internally and was critical for successful AI adoption. For instance, TSE recognized the value of bringing plant operators into the ML model development and began holding “office hours” in which they and project managers could discuss the analytics models with data scientists. Data scientists at Zymergen also found that upfront and frequent involvement of physical scientists in the machine learning model was critical to building trust for successful implementation.

Across the three organizations, explainability of AI systems played a critical role in building this trust with employees, customers, and regulators.

Each firm reported productivity or financial gains in the short term.

TSE reports it experienced productivity gains through improved production yield, reduced raw material expenses, and enhanced product quality. Axis Bank reports it was able to handle its growing customer service volumes with fewer customer service agents, driven by migration from human-enabled to  automated customer service channels.

Lastly, in the Zymergen case, management reported higher labor productivity, driven by a high degree of automation in the wet lab and accelerated project durations compared with conventional R&D labs.

In all cases, however,  these effects could not be attributed solely to AI, as other process changes always accompanied the introduction of AI technologies.

Workforce impacts varied and in some  cases, cascaded beyond the firms.

In the three cases, we observed varying degrees of labor impacts, both direct and indirect, demonstrating that labor effects are often multi-layered and can extend beyond the core organization implementing AI. The case studies also demonstrated that industry and regulatory contexts play a role in how AI-related initiatives impact labor.

For instance, Tata Steel Europe operates in a mature and highly unionized industry with a relatively inflexible labor market. As a result, TSE specifically avoided workforce reduction, although this may not be the case in the long-term.

At Axis Bank, the effects of AI implementation most directly impacted its third-party customer service provider rather than internal employees. Similarly, Zymergen’s business model may have external, downstream impacts on its customers’ R&D teams, potentially leading to lower hiring rates in the future. As these cases demonstrate, the social and economic impact of AI and other automating technologies goes beyond the immediate sites of implementation and cascades across supply chains, partners, customers, and others affected by the technology.

Though we observed certain common themes across the three case studies, each organization faced its own context-specific opportunities and challenges. For instance, we note greater flexibility for Zymergen – an “AI-native” company – in pursuing its programs, than for longer-established companies, particularly those that may have greater regulatory complexity or unionized workforces.

Scholars, managers, policymakers, and others should not overlook these contextual factors when trying to gauge AI’s impact on individual organizations and the broader economy.

Lastly, we acknowledge that this group of case studies is not comprehensive and that its chosen methodology and scope has limitations. For instance, our interviews were primarily with management at each of the subject organizations. One valuable area for future research would be collecting and synthesizing data directly from workers directly impacted by AI implementation. As documentation of the actual impact of real-world AI applications, albeit partial, the case studies can contribute nuanced pictures of the way that AI and ML are playing out in real workplaces. These observations and insights should inform ongoing dialogue on the impact of AI on labor and the economy.

Learn more at the AI Case Study Compendium.

Peter Lo is Senior Communications Manager at Partnership on AI. He can be reached at Peter.lo@partnershiponai.org.

How Machine Learning Is Reshaping Location-Based Services

When we interact with digital devices and services, they almost always generate data about our location. Smartphone users rely on location services to look up driving directions and weather reports and set up geo-fenced alerts. For service providers, this information is equally useful: It can reveal a lot about customers, competitors, and opportunities to expand or improve their services.

Applying machine learning will make location-based services even better attuned to our needs and preferences. Let’s take a look at what this might mean in practice.

Proactive Nagivation

The greatest triumph of the smartphone was the democratization of the Global Positioning System. Once just a tool for governments and militaries, GPS now empowers people all over the globe with insights into where they’re going and how to get there. We can complain all we like that nobody knows how to read a paper map anymore, but does anybody really want to go back to those days?

Thanks to machine learning, our smartphones, and mobile apps are going to get better at delivering uncannily accurate predictions about where we need to go, when we need to leave and how to get there, all based on pattern recognition and historical user data. We can already add a “time to leave” modifier to the entries we make in our calendars, but machine learning will take this to the next level.

Next time you’re getting ready for work, imagine receiving an unprompted notification on your phone screen indicating a traffic snarl-up along your usual route. You know how to get to the office — you haven’t needed directions to get there in a long time. However, thanks to machine learning, your phone is helpfully pinging you with an alternate route, because it knows your usual commute is going to slow you down and it doesn’t want you to be late.

From Apple and Google to Nokia and lots of startups you haven’t even heard of yet, there’s a lot of money being poured into intelligent navigation systems. The future, according to Uber, Lyft, Tesla and others, is autonomous cars with smart navigation that can change in response to real-time events. Getting there means our technology needs to get a lot better at studying and visualizing user patterns for thousands of customers at one time. It also needs to take into account congestion, weather, time of day, planned and unplanned events, and much more.

Smarter Apps

Many of us rely on smartphones to remind us about upcoming items on our to-do lists, to keep track of which groceries we’re running low on, and whether our next dental cleaning is coming up. If reminders and calendar items are the bread and butter of the mobile operating system experience, machine learning is the secret sauce that could take stock smartphone apps and deliver performance that makes it feel like we’re living in the future.

Suppose we’ve been adding items to our grocery lists like toilet paper, milk, and eggs. Before too long, our favorite apps will be smart enough to plan our shopping trips and even entire days with uncanny accuracy and help us make the most of our limited time. They’ll know what’s on our shopping or to-do lists and where we’ve been in the past when we checked those items off. The next time we’re driving by the store, they’ll let us know about it — all without being asked — or maybe even suggest an alternative that’s running a sale or promotion.

We tell ourselves that smartphones are like digital personal assistants, but we have to perform a lot of the logic for them. Thanks to a combination of machine learning and location services, we can expect far more intuitive and automated performance soon.

Read the source article in TechiExpert.com.

Blanket ban on cryptocurrency may prove to be futile, claim experts

Experts believe the decentralised nature of cryptocurrency may render such a ban futile

Seven Guidelines to Ensure Ethical AI

The organisation of tomorrow will be built around data, and it will require artificial intelligence to make sense of all that data. Artificial intelligence is a broad discipline with the objective to develop intelligent machines.

AI consists of several subfields: Machine learning (ML), a subset of AI that enables machines to learn from data. Reinforcement learning, which is a subset of ML and focuses on artificial agents that use trial and error to improve itself. And deep learning, also a subset of ML that aims to mimic the human brain to detect patterns in large datasets and benefit from those patterns.

Artificial intelligence has been around since the 1950s, thanks to the work of Alan Turing, who is widely regarded as the father of theoretical computer science and artificial intelligence. Back then, Alan Turing already dreamed of machines that would eventually “compete with men in all purely intellectual fields”.

In the past years, we have come to a lot closer to Turing's dream. Thanks to billions of dollars in research spent by tech giants such as Google, Facebook, Tencent, Baidu, Apple and Microsoft, there has been increased attention on AI. Resulting in a variety of ever-more intelligent AI applications in almost every domain.

The ...


Read More on Datafloq

What it Really Takes to Become a Professional Data Scientist

Data science, is it a piece of cake? A data scientist is still deemed as the sexiest job of the 21st century. Though the profession had been ranked on the top of the job lists three years in a row, the deficit for such professionals is on a constant rise.

As estimated by the European Commission, it is projected that 100,000 new data jobs will be available by 2020. While the truth still prevails, will there be enough supply for data science professionals? The organizations now have awakened to the fear that the demand will outpace the supply. With businesses on the rise, it is evident that data science has the potential to drive organizations to a different level altogether. IBM projects that the data science jobs will account 28% of all other digital jobs by 2020. On another report, it is stated that on an average, most of the unfilled data science jobs remained vacant for nearly 45 days, since those who were applying for the job role do not have the relevant skills to become a data science professional.

Over the past years, technology skills such as machine learning, big data, and data science have given birth to plenty of ...


Read More on Datafloq

How Data Science is Transforming Web Development

“A billion hours ago, modern Homo sapiens emerged.
A billion minutes ago, Christianity began.
A billion seconds ago, the IBM personal computer was released.
A billion Google searches ago…was this morning."
-Hal Varian, Google’s Chief Economist, December, 2013 (From the book: Work Rules by Laszlo Buck)


The last line of the above quote points at the world’s hunger for information. Information plays a huge role in our life. Information consumed by our senses helps our mind in taking decisions. But what happens when the mind is flooded with information. You get confused, irritated and scared of decision-making. This is where your computers and processors come to rescue, and this is when the term “information” is replaced by “data.”

Every minute, more than a hundred hours of video content is uploaded on YouTube. From the application stores, 50+ billion apps have already been downloaded since 2008. There are 2+ billion people signed up on social media websites. These numbers are just giving you a glimpse of the amount of data which is flowing through the optical fibers every second around the world. And now the question comes “how to make this massive amount of data useful?” The answer is Analytics. If you know how to play with ...


Read More on Datafloq

Inside India's push to build an indigenous semiconductor design ecosystem

Several local companies and academia-industry incubators have mushroomed across the country, designing and testing chipsets, microprocessors and allied technology that could be used commercially.

Scooter-borne first responders could prove vital for emergency healthcare

Heart Rescue India, a partnership between Ramaiah Memorial Hospital and the US-based University of Illinois, has taken the help of Bengaluru-based AI startup Inkers.ai to put together such an emergency response protocol.

Thursday, 2 May 2019

Detecting Financial Fraud at Scale with Decision Trees and MLflow on  Databricks

Try this notebook in Databricks

Detecting fraudulent patterns at scale is a challenge, no matter the use case. The massive amounts of data to sift through, the complexity of the constantly evolving techniques, and the very small number of actual examples of fraudulent behavior are comparable to finding a needle in a haystack while not knowing what the needle looks like. In the world of finance, the added concerns with security and the importance of explaining how fraudulent behavior was identified further increases the complexity of the task.

To build these detection patterns, a team of domain experts often comes up with a set of rules that define fraudulent behavior. A typical workflow may include a subject matter expert in the financial fraud detection space putting together a set of requirements for a particular behavior. A data scientist may then take a subsample of the available data and build a model using these requirements and possibly some known fraud cases. To put the pattern in production, a data engineer may convert the resulting model to a set of rules with thresholds, often implemented using SQL.

This approach allows the financial institution to present a clear set of characteristics that led to the identification of fraud that is compliant with the General Data Protection Regulation (GDPR). However, this approach also poses numerous difficulties. The implementation of the detection pattern using a hardcoded set of rules is very brittle. Any changes to the pattern would take a very long time to update. This, in turn, makes it difficult to keep up with and adapt to the shift in fraudulent behaviors that are happening in the current marketplace.

Additionally, the systems in the workflow described above are often siloed, with the domain experts, data scientists, and data engineers all compartmentalized. The data engineer is responsible for maintaining massive amounts of data and translating the work of the domain experts and data scientists into production level code. Due to a lack of common platform, the domain experts and data scientists have to rely on sampled down data that fits on a single machine for analysis. This leads to difficulty in communication and ultimately a lack of collaboration.

In this blog, we will showcase how to convert several such rule-based detection use cases to machine learning use cases on the Databricks platform, unifying the key players in fraud detection: domain experts, data scientists, and data engineers. We will learn how to create a fraud-detection data pipeline and visualize the data leveraging a framework for building modular features from large data sets. We will also learn how to detect fraud using decision trees and Apache Spark MLlib. We will then use MLflow to iterate and refine the model to improve its accuracy.

Solving with ML

There is a certain degree of reluctance with regard to machine learning models in the financial world as they are believed to offer a “black box” solution with no way of justifying the identified fraudulent cases. GDPR requirements, as well as financial regulations, make it seemingly impossible to leverage the power of machine learning. However, several successful use cases have shown that applying machine learning to detect fraud at scale can solve a host of the issues mentioned above.

 

Training a supervised machine learning model to detect financial fraud is very difficult due to the low number of actual confirmed examples of fraudulent behavior. However, the presence of a known set of rules that identify a particular type of fraud can help create a set of synthetic labels and an initial set of features. The output of the detection pattern that has been developed by the domain experts in the field has likely gone through the appropriate approval process to be put in production. It produces the expected fraudulent behavior flags and may, therefore, be used as a starting point to train a machine learning model. This simultaneously mitigates three concerns:

  1. The lack of training labels,
  2. The decision of what features to use,  and
  3. Having an appropriate benchmark for the model.

Training a machine learning model to recognize the rule-based fraudulent behavior flags offers a direct comparison with the expected output via a confusion matrix. Provided that the results closely match the rule-based detection pattern, this approach helps gain confidence in machine learning based fraud detection with the skeptics. The output of this model is very easy to interpret and may serve as a baseline discussion of the expected false negatives and false positives when compared to the original detection pattern.

Furthermore, the concern with machine learning models being difficult to interpret may be further assuaged if a decision tree model is used as the initial machine learning model. Because the model is being trained to a set of rules, the decision tree is likely to outperform any other machine learning model. The additional benefit is, of course, the utmost transparency of the model, which will essentially show the decision-making process for fraud, but without human intervention and the need to hard code any rules or thresholds. Of course, it must be understood that the future iterations of the model may utilize a different algorithm altogether to achieve maximum accuracy. The transparency of the model is ultimately achieved by understanding the features that went into the algorithm. Having interpretable features will yield interpretable and defensible model results.

The biggest benefit of the machine learning approach is that after the initial modeling effort, future iterations are modular and updating the set of labels, features, or model type is very easy and seamless, reducing the time to production. This is further facilitated on the Databricks Unified Analytics Platform where the domain experts, data scientists, data engineers may work off the same data set at scale and collaborate directly in the notebook environment. So let’s get started!

Ingesting and Exploring the Data

We will use a synthetic dataset for this example. To load the dataset yourself, please download it to your local machine from Kaggle and then import the data via Import Data – Azure and AWS

The PaySim data simulates mobile money transactions based on a sample of real transactions extracted from one month of financial logs from a mobile money service implemented in an African country. The below table shows the information that the data set provides:

Exploring the Data

Creating the DataFrames – Now that we have uploaded the data to Databricks File System (DBFS), we can quickly and easily create DataFrames using Spark SQL


# Create df DataFrame which contains our simulated financial fraud detection dataset
df = spark.sql("select step, type, amount, nameOrig, oldbalanceOrg, newbalanceOrig, nameDest, oldbalanceDest, newbalanceDest from sim_fin_fraud_detection")

Now that we have created the DataFrame, let’s take a look at the schema and the first thousand rows to review the data.


# Review the schema of your data
df.printSchema()
root
|-- step: integer (nullable = true)
|-- type: string (nullable = true)
|-- amount: double (nullable = true)
|-- nameOrig: string (nullable = true)
|-- oldbalanceOrg: double (nullable = true)
|-- newbalanceOrig: double (nullable = true)
|-- nameDest: string (nullable = true)
|-- oldbalanceDest: double (nullable = true)
|-- newbalanceDest: double (nullable = true)

Types of Transactions

Let’s visualize the data to understand the types of transactions the data captures and their contribution to the overall transaction volume.

To get an idea of how much money we are talking about, let’s also visualize the data based on the types of transactions and on their contribution to the amount of cash transferred (i.e. sum(amount)).

Rules-based Model

We are not likely to start with a large data set of known fraud cases to train our model. In most practical applications, fraudulent detection patterns are identified by a set of rules established by the domain experts. Here, we create a column called label based on these rules.


# Rules to Identify Known Fraud-based
df = df.withColumn("label", 
                   F.when(
                     (
                       (df.oldbalanceOrg <= 56900) & (df.type == "TRANSFER") & (df.newbalanceDest <= 105)) | ( (df.oldbalanceOrg > 56900) & (df.newbalanceOrig <= 12)) | ( (df.oldbalanceOrg > 56900) & (df.newbalanceOrig > 12) & (df.amount > 1160000)
                           ), 1
                   ).otherwise(0))

Visualizing Data Flagged by Rules

These rules often flag quite a large number of fraudulent cases. Let’s visualize the number of flagged transactions. We can see that the rules flag about 4% of the cases and 11% of the total dollar amount as fraudulent.

Selecting the Appropriate Machine Learning Models

In many cases, a black box approach to fraud detection cannot be used. First, the domain experts need to be able to understand why a transaction was identified as fraudulent. Then, if action is to be taken, the evidence has to be presented in court. The decision tree is an easily interpretable model and is a great starting point for this use case. Read this blog “The wise old tree” on decision trees to learn more.

Creating the Training Set

To build and validate our ML model, we will do an 80/20 split using .randomSplit. This will set aside a randomly chosen 80% of the data for training and the remaining 20% to validate the results.


# Split our dataset between training and test datasets
(train, test) = df.randomSplit([0.8, 0.2], seed=12345)

Creating the ML Model Pipeline

To prepare the data for the model, we must first convert categorical variables to numeric using .StringIndexer. We then must assemble all of the features we would like for the model to use. We create a pipeline to contain these feature preparation steps in addition to the decision tree model so that we may repeat these steps on different data sets. Note that we fit the pipeline to our training data first and will then use it to transform our test data in a later step.


from pyspark.ml import Pipeline
from pyspark.ml.feature import StringIndexer
from pyspark.ml.feature import VectorAssembler
from pyspark.ml.classification import DecisionTreeClassifier

# Encodes a string column of labels to a column of label indices
indexer = StringIndexer(inputCol = "type", outputCol = "typeIndexed")

# VectorAssembler is a transformer that combines a given list of columns into a single vector column
va = VectorAssembler(inputCols = ["typeIndexed", "amount", "oldbalanceOrg", "newbalanceOrig", "oldbalanceDest", "newbalanceDest", "orgDiff", "destDiff"], outputCol = "features")

# Using the DecisionTree classifier model
dt = DecisionTreeClassifier(labelCol = "label", featuresCol = "features", seed = 54321, maxDepth = 5)

# Create our pipeline stages
pipeline = Pipeline(stages=[indexer, va, dt])

# View the Decision Tree model (prior to CrossValidator)
dt_model = pipeline.fit(train)

Visualizing the Model

Calling display() on the last stage of the pipeline, which is the decision tree model, allows us to view the initial fitted model with the chosen decisions at each node. This helps to understand how the algorithm arrived at the resulting predictions.


display(dt_model.stages[-1])

Visual representation of the Decision Tree model

Model Tuning

To ensure we have the best fitting tree model, we will cross-validate the model with several parameter variations. Given that our data consists of 96% negative and 4% positive cases, we will use the Precision-Recall (PR) evaluation metric to account for the unbalanced distribution.


from pyspark.ml.tuning import CrossValidator, ParamGridBuilder

# Build the grid of different parameters
paramGrid = ParamGridBuilder() \
.addGrid(dt.maxDepth, [5, 10, 15]) \
.addGrid(dt.maxBins, [10, 20, 30]) \
.build()

# Build out the cross validation
crossval = CrossValidator(estimator = dt,
                          estimatorParamMaps = paramGrid,
                          evaluator = evaluatorPR,
                          numFolds = 3)  
# Build the CV pipeline
pipelineCV = Pipeline(stages=[indexer, va, crossval])

# Train the model using the pipeline, parameter grid, and preceding BinaryClassificationEvaluator
cvModel_u = pipelineCV.fit(train)

 

Model Performance

We evaluate the model by comparing the Precision-Recall (PR) and Area under the ROC curve (AUC) metrics for the training and test sets. Both PR and AUC appear to be very high.


# Build the best model (training and test datasets)
train_pred = cvModel_u.transform(train)
test_pred = cvModel_u.transform(test)

# Evaluate the model on training datasets
pr_train = evaluatorPR.evaluate(train_pred)
auc_train = evaluatorAUC.evaluate(train_pred)

# Evaluate the model on test datasets
pr_test = evaluatorPR.evaluate(test_pred)
auc_test = evaluatorAUC.evaluate(test_pred)

# Print out the PR and AUC values
print("PR train:", pr_train)
print("AUC train:", auc_train)
print("PR test:", pr_test)
print("AUC test:", auc_test)

---
# Output:
# PR train: 0.9537894984523128
# AUC train: 0.998647996459481
# PR test: 0.9539170535377599
# AUC test: 0.9984378183482442

To see how the model misclassified the results, let’s use matplotlib and pandas to visualize our confusion matrix.

 

Balancing the Classes

We see that the model is identifying 2421 more cases than the original rules identified. This is not as alarming as detecting more potential fraudulent cases could be a good thing. However, there are 58 cases that were not detected by the algorithm but were originally identified. We are going to attempt to improve our prediction further by balancing our classes using undersampling.  That is, we will keep all the fraud cases and then downsample the non-fraud cases to match that number to get a balanced data set. When we visualized our new data set, we see that the yes and no cases are 50/50.


# Reset the DataFrames for no fraud (`dfn`) and fraud (`dfy`)
dfn = train.filter(train.label == 0)
dfy = train.filter(train.label == 1)

# Calculate summary metrics
N = train.count()
y = dfy.count()
p = y/N

# Create a more balanced training dataset
train_b = dfn.sample(False, p, seed = 92285).union(dfy)

# Print out metrics
print("Total count: %s, Fraud cases count: %s, Proportion of fraud cases: %s" % (N, y, p))
print("Balanced training dataset count: %s" % train_b.count())

---
# Output:
# Total count: 5090394, Fraud cases count: 204865, Proportion of fraud cases: 0.040245411258932016
# Balanced training dataset count: 401898
---

# Display our more balanced training dataset
display(train_b.groupBy("label").count())

Updating the Pipeline

Now let’s update the ML pipeline and create a new cross validator. Because we are using ML pipelines, we only need to update it with the new dataset and we can quickly repeat the same pipeline steps.


# Re-run the same ML pipeline (including parameters grid)
crossval_b = CrossValidator(estimator = dt,
estimatorParamMaps = paramGrid,
evaluator = evaluatorAUC,
numFolds = 3)
pipelineCV_b = Pipeline(stages=[indexer, va, crossval_b])

# Train the model using the pipeline, parameter grid, and BinaryClassificationEvaluator using the `train_b` dataset
cvModel_b = pipelineCV_b.fit(train_b)

# Build the best model (balanced training and full test datasets)
train_pred_b = cvModel_b.transform(train_b)
test_pred_b = cvModel_b.transform(test)

# Evaluate the model on the balanced training datasets
pr_train_b = evaluatorPR.evaluate(train_pred_b)
auc_train_b = evaluatorAUC.evaluate(train_pred_b)

# Evaluate the model on full test datasets
pr_test_b = evaluatorPR.evaluate(test_pred_b)
auc_test_b = evaluatorAUC.evaluate(test_pred_b)

# Print out the PR and AUC values
print("PR train:", pr_train_b)
print("AUC train:", auc_train_b)
print("PR test:", pr_test_b)
print("AUC test:", auc_test_b)

---
# Output: 
# PR train: 0.999629161563572
# AUC train: 0.9998071389056655
# PR test: 0.9904709171789063
# AUC test: 0.9997903902204509

Review the Results

Now let’s look at the results of our new confusion matrix. The model misidentified only one fraudulent case. Balancing the classes seems to have improved the model.

Model Feedback and Using MLflow

Once a model is chosen for production, we want to continuously collect feedback to ensure that the model is still identifying the behavior of interest. Since we are starting with a rule-based label, we want to supply future models with verified true labels based on human feedback. This stage is crucial for maintaining confidence and trust in the machine learning process. Since analysts are not able to review every single case, we want to ensure we are presenting them with carefully chosen cases to validate the model output. For example, predictions, where the model has low certainty, are good candidates for analysts to review. The addition of this type of feedback will ensure the models will continue to improve and evolve with the changing landscape.

MLflow helps us throughout this cycle as we train different model versions. We can keep track of our experiments, comparing the results of different model configurations and parameters. For example here, we can compare the PR and AUC of the models trained on balanced and unbalanced data sets using the MLflow UI. Data scientists can use MLflow to keep track of the various model metrics and any additional visualizations and artifacts to help make the decision of which model should be deployed in production. The data engineers will then be able to easily retrieve the chosen model along with the library versions used for training as a .jar file to be deployed on new data in production. Thus, the collaboration between the domain experts who review the model results, the data scientists who update the models, and the data engineers who deploy the models in production, will be strengthened throughout this iterative process.

Conclusion

We have reviewed an example of how to use a rule-based fraud detection label and convert it to a machine learning model using Databricks with MLflow. This approach allows us to build a scalable, modular solution that will help us keep up with ever-changing fraudulent behavior patterns. Building a machine learning model to identify fraud allows us to create a feedback loop that allows the model to evolve and identify new potential fraudulent patterns. We have seen how a decision tree model, in particular, is a great starting point to introduce machine learning to a fraud detection program due to its interpretability and excellent accuracy.

A major benefit of using the Databricks platform for this effort is that it allows for data scientists, engineers, and business users to seamlessly work together throughout the process. Preparing the data, building models, sharing the results, and putting the models into production can now happen on the same platform, allowing for unprecedented collaboration. This approach builds trust across the previously siloed teams, leading to an effective and dynamic fraud detection program.

Try this notebook by signing up for a free trial in just a few minutes and get started creating your own models.

--

Try Databricks for free. Get started today.

The post Detecting Financial Fraud at Scale with Decision Trees and MLflow on  Databricks appeared first on Databricks.

Is Ai a Good-fit For Small Businesses? Or Dp We Need Human Intelligence Too?

Artificial Intelligence is getting an upswing day-by-day. No doubt, it is a realistic and emerging innovation but there also present limitations that businesses must aware of. AI is a complex solution that is more expensive for small businesses due to their limited data and resources.

Dan Faggella, a renowned AI expert, said something for this,

“Almost 99.9% small businesses don’t even need AI to earn profitable outcomes right now.”

Cost-Effective For Small Businesses?

As machines need a lubricant to work properly, algorithms behind an AI also needs to be optimized and changed with respect to the operational conditions. For regular AI updates, investment is required which is a big hassle for small businesses which is unavoidable.

Should You Rely On AI Only!

Do you think, AI can’t go wrong? AI works on algorithms, it can’t resolve issues or concerns on its own. It definitely needs a human touch to work properly. It can go wrong at any stage. It should be tested precisely before letting it go to work on your brand.

You should monitor and track when you need to modify its algorithms. In this regards, as a small business owner, you would suffer because working on AI with limited resources is not completely beneficial. AI ...


Read More on Datafloq

Why C-Level Executives Need to Close the Digital Transformation Disparity Gap

The digital transformation has become such an important part of organizational culture today that it is extremely easy for stakeholders to forget what it actually means. The fact that this term has been thrown around quite a lot recently has meant that we now have a disparity in what stakeholders perceive about this term and the actual meaning that it carries. 

This disparity in perception and the actual progress is concerning. C-level executives believe that they have made solid progress towards digital implementation, but the reality is often starkly different. Here, we shed some light onto the digital transformation, and how C-level executives should be on the same page as everyone else in identifying the progress made towards a data-oriented future for their organization. 

I am really excited to let you know that this article has been written in partnership with Tata Communications based on their ‘Cycle of Progress in Digital Transformation’ research. 

Understanding the Perception vs. Reality Gap 

While many businesses in this day and age have taken up the steps to implement AI within their organization, there is still a significant disparity with the level of progress C-level executives believe they have made in this area versus reality.

To begin their research, Tata Communications interviewed stakeholders ...


Read More on Datafloq

India working on robots that may patrol borders

As part of enhancing India's defence capabilities, scientists have been quietly working on all-terrain artificial intelligence (AI)-enabled robots that may eventually patrol the country's borders.

Wednesday, 1 May 2019

5 Ways AI and Big Data Are Revolutionizing Education

While the world goes smart at an astonishing pace, turning phones into personal health monitors and TVs into voice-controlled streaming devices; educational institutions cannot afford to remain archaic stone buildings with rigid curriculums and one-dimensional grading systems. Artificial intelligence and big data are helping schools, colleges and universities become more sophisticated and better capable of helping more students attain a better quality of education. One that will better enable them to attain their highest potential and become valuable members of society.

From delivering highly engaging lectures that are better understood by students, to performing more intuitive aptitude assessments to propel students into the right courses for higher education, AI and big data are helping change the very course of formal education and bringing it closer to the goal it was originally intended for – to inform young minds and enable them to make the best use of that information.

So let’s take a look at all the ways AI and big data solve the biggest problems students and institutions face.

1. Improved Effectiveness of Learning through Personalized Program

One of the major players in a conventional classroom setting is student diversity. Some students are naturally good at grasping mathematical concepts while others struggle with ...


Read More on Datafloq

Genderless Voice AI Could Provide Major Step in Combating Implicit Bias

Implicit in any technical process or system are the biases of those writing the code that will govern the actions of that respective technical system or process. I’m not throwing shade at developers in saying that, but rather highlighting that we all suffer from implicit biases — whether known or not — and those biases get baked into the software solutions we develop and deliver. 

We’ve written about it before, but I think it bears repeating because there are some pretty fascinating solutions on deck aimed at combatting some common but probably unrecognized variants of this. Namely, the interface of the future, natural language processing (NLP), is confined to binary voice characteristics unnecessarily. Enter the genderless voice AI ‘Q’.

Why do our vocal assistants’ voices matter?

Quoting a great paragraph from Mark Wilson in Fast Company:

“Voice assistants like Apple’s Siri and Amazon’s Alexa are women rather than men. You can change this in the settings, and choose a male speaker, of course, but the fact that the technology industry has chosen a woman to, by default, be our always-on-demand, personal assistant of choice, speaks volumes about our assumptions as a society: Women are expected to carry the psychic burden of schedules, birthdays, and phone numbers; they ...


Read More on Datafloq