Thursday, 26 September 2019

14 Ways Machine Learning Can Boost Your Marketing Introduction

Machine learning is an application of artificial intelligence. And it is no secret that this technology is revolutionizing the marketing niche. It enhances a system in a way it that it learns from current and past experiences. It involves the invention and development of computer programs that can read data and use it to influence performance. This, of course, happens without human intervention as the whole process is automated. Exploring technology for gains in business is a practice that has taken ground. More and more companies are employing artificial intelligence in their daily operations, in a bid to perfect customer experience and boost returns. Thanks to the rapid growth in technology, marketing has seen different improvements.

So, how can machine learning improve your business and affect the flow of clients? This is an argument that has elicited mixed reactions from pundits in the industry. Some argue that this is the way to go for the future, while others are a bit skeptical and reserved. While conventional means of advertising worked efficiently, this may not be the trend throughout. Change is inevitable, and so, you need to shape yourself ready for it, if you’re a participant in this niche. Besides, with the ...


Read More on Datafloq

How Predictive Analytics is Helping Organizations to Increase Sales

Predictive analytics is a process that involves passing historical data through statistical algorithms to identify the likelihood of future events. The technology has been around for decades now, with various statistical models in use as early as the 80s to predict the stock market.

Only recently, however, have the powers that be come together and made an undeniably strong case for wider adoption of the once unstable technology.

What is Driving Predictive Analytics in Sales?

Growing volumes of data, in part due to more ways to collect them, faster computers for a fraction of the price and cloud computing have all come together to create the perfect environment for predictive analytics.

Even as new technology such as machine learning, artificial intelligence and in-memory computing continue to take root, the impact on the internet economy is expected to be massive. In particular, the latter of these (in-memory computing) has spearheaded the availability of real-time data, greatly increasing the speed with which analytics can be utilized.

In-memory computing is basically the processing of large volumes of data in system memory rather than relying on disk-based processing such as transactional databases. The most impactful implication of this is the resulting speed and performance boost it can achieve.

Few places are ...


Read More on Datafloq

JavaFX: What Makes it Ideal for Cross-Platform App Development Projects

App development is an inherently complex endeavour from the get-go, necessitating the close and thoughtful analysis of myriad factors and demanding many, many choices before and during the project. One such question any company or team planning to develop a cross-platform app is the library they will use. Now, this is a crucial question because it plays a critical role in deciding just how well the app’s UX turns out. Also, as any developer will tell you, the UX is the deciding factor when it comes to the app’s potential to succeed and make an impact.

Much like anything else, there are plenty of options in the market in this regard as well, but we’ll focus our attention on just one leading name: JavaFX. The list of reasons why this name, in particular, stands out from the crowd is long! For example, JavaFX makes use of a high-quality graphics pipeline to deliver intricate UI rendering that functions exceptionally well. It is especially important when one is looking to put together a high-performance app or perhaps a 3D one. However, to truly help you understand this tool’s potential, allow us to walk you through an in-depth list of the reasons that make ...


Read More on Datafloq

Army plans to induct AI to bolster capability

These mechanised formations are largely deployed along India’s western front with Pakistan and in some sectors of the northern border with China such as Ladakh and North Sikkim.

Wednesday, 25 September 2019

Making the Move to Amsterdam: Miles Yucht

While we are proud of our Berkeley roots, Databricks now calls many cities around the world our home. In addition to offices in London, Singapore, New York and our headquarters in San Francisco, we have one of our major engineering hubs in the fast-growing European Development Center, Amsterdam. Databricks offers the exciting opportunity to relocate to one of our global offices and we are proud to say that we have been able to help our employees navigate through that transition. Learn more about Miles Yucht, a Tech Lead on the Billing Team, and what inspired him to make the move from San Francisco to Amsterdam.

Miles Yucht(on the right) with the Databricks Amsterdam team at a post-apocalypse themed escape room

Miles (on the right) with the team at a post-apocalypse themed escape room

Tell us a little about yourself.

I studied computer science and music at Princeton, and moved to California afterwards, where I joined Databricks and helped scale our San Francisco office. A year ago, I got the opportunity to work in Amsterdam for two weeks, where I got to know the local team, explore the city and get a better feel of what it’d be like to live there. In March 2019, I decided to move to Amsterdam, where I lead the Billing Infrastructure and Usage team. Our team develops process and solutions to bill customers appropriately while giving them visibility on how to use the product efficiently.

What were you looking for in your next opportunity?

I really wanted to experience living somewhere completely different. I grew up in the remote countryside in Vermont, and lived in one of the biggest cities in the U.S., but hadn’t gotten the opportunity to see other parts of the world that I wanted to, including Europe. I was looking for cultural exchange, and the opportunity to meet other people around the world and understand their perspectives. The opportunity to move and work with Databricks abroad in Amsterdam and be a part of building out the European Development Center was very exciting. It also offered an opportunity to learn about the Dutch viewpoint and lifestyle.

How did you choose to go to Amsterdam and Databricks specifically?

Amsterdam at Dusk

Amsterdam at Dusk

In college, I traveled with my school orchestra for an international tour in Germany and the Netherlands. When I was in Amsterdam, I fell completely in love with the place – seeing the appreciation for biking, how safe it was, the love for classical music and art, and the beautiful scenery, including all of the canals and brick houses. Amsterdam was a major city, but still a little more quaint than the hustle-bustle of usual cities, which really connected with me as someone who came from a more rustic background.

That being said, when I was looking for computer science careers as I was graduating, I was focused more on jobs located in the Bay Area. My roommate had met Patrick Wendell at an alumni event, who convinced him to intern at the company he was founding, Databricks. We both applied and received internships there, however, I decided to intern at a larger company, while he interned at Databricks. I had a lot of fun at my internship (there were a lot of fun events, a big nice office), but I didn’t feel that I’d be able to push myself in terms of my own personal growth, versus the opportunity that Databricks offered (and that he had experienced), which at the time was only a 50-person company. When I graduated, I decided to join Databricks full-time. Fast forward to December 2016, Databricks had hired a core group of Database engineers in Amsterdam and opened an office there. This was the moment that I realized that my dream of moving to Amsterdam could actually happen and that I’d potentially be able to transfer in the future. Two years later, it finally happened!

What was the biggest challenge you faced when relocating to Amsterdam, and what lessons did you learn from it?

I didn’t realize how many basic things I took for granted that I needed to sort out – like getting a new SIM card, getting my visa, and situating my bank account. Luckily, the Dutch government does a great job of making the transition smooth – they provide a lot of helpful resources on what to do on the first day, month, and even have a SIM card you can use in your first few weeks here. Adjusting to the Dutch environment, learning a new language, and finding a local friend group, was also challenging at times. However, most people here are fluent in English and really nice. I was lucky to meet a group of local students and musicians through joining music groups. It definitely wasn’t easy, but it taught me to welcome the opportunity to pursue extracurriculars when moving to a new city. It’s a great way to meet people outside of work and to experience the city through multiple perspectives.

What has been the most exciting aspect about living in Amsterdam?

Working abroad with Databricks: Miles Yucht enjoys yoga at the Amsterdam Rijksmuseum with coworker

Yoga at the Rijksmuseum with my coworker

It’s been very exciting to see the level of pride and passion that Dutch people have for their country – from sports to art! There is also so much culture in Amsterdam. For instance, there are so many museums, music groups, and stunning venues that are easily accessible. You can even utilize a museum pass to visit many museums and do fun events such as yoga at the Rijksmuseum! It’s also exciting to be able to travel to other places so easily. I’ve already been to France and Italy and am planning on visiting Germany and Austria in the next few weeks!

I’ve also gotten to really know the Amsterdam team, that I previously was only able to meet a couple weeks at a time when I was working in our SF office. For me, it’s exciting to be part of the first team from our Platform division to be based in Amsterdam, serve as a vanguard for our department, all while helping develop the cultural building blocks for our rapidly growing office! You don’t get that opportunity very often.

Databricks has grown tremendously in the last few years. How do you see the future of Databricks evolving and what are you most excited to see us accomplish?

We’ve been hiring as quickly and sustainably as we can, so we can provide the highest amount of value to all our customers. I’ve noticed that even though our headquarters are in San Francisco, our global hiring has skyrocketed – all of our global offices are looking to expand the office space they’re staying in. Seeing all that growth makes me think that we’ll be a lot more globally distributed in the future. Who knows, maybe there will be a company retreat in Amsterdam one day!

In addition, the direction that we’ve taken to be the home of Machine Learning and Big Data is a really important decision that our leaders made from day one. The more I see our success, the more I’m convinced it’s the right answer. I’m really excited to see Databricks become the dominant player for Machine Learning and Big Data, and for everyone to want to use our unified analytics platform because we have an amazing product!

What advice would you give to people considering relocation?

If you have the opportunity, you should grab the chance to live in a different part of the world! It really broadens your world view around all types of issues, from healthcare to immigration, and gives you perspective on your own life. Being able to see different problems people face in their day-to-day, and how they solve it, teaches you a lot about how to work with different people and expands your perspective. Also, even if you make great friends abroad, don’t lose touch with the people you moved away from. It’s easy to not message someone for a couple of weeks, especially if you’re used to talking to people in-person versus messaging. However, it’s worth investing enough time to maintain relationships you had wherever you’re coming from.

Interested in joining Miles’ team in our Amsterdam office? Check out our Careers Page.

--

Try Databricks for free. Get started today.

The post Making the Move to Amsterdam: Miles Yucht appeared first on Databricks.

2019 Jenkins Board and Officer elections. Nominations are open!

This is a repost of the original announcement made by Kohsuke Kawaguchi in the Jenkins Developer mailing list. Minor changes were applied to reflect the posting date and to provide more links.

Nominations for the 2019 Jenkins Board elections open for three governing board positions and five officer positions, namely: Security, Events, Release, Infrastructure and Documentation.

The terms of office for these positions are:

  • Officer positions (1 year): November 4, 2019 to November 3, 2020

  • Governing board members (2 years): November 4, 2019 to November 3, 2021

To nominate someone, simply send an email to jenkinsci-board@googlegroups.com with their name and position you nominate them for. Please share any information on why you are making the nomination. Self nominations are also welcome.

The board positions and officer roles are an essential part of Jenkins' community governance and well-being. I highly encourage everyone to consider participating.

Key dates

  • Oct 04, 2019: Nominations close

  • Oct 08, 2019: List of nominees posted to mailing list

  • Oct 11, 2019: Nominees’ personal statements made available

  • Oct 14, 2019: Voting begins

  • Oct 27, 2019: Voting closes at 5pm Pacific Time

  • Nov 04, 2019: New representatives announced

A Guide to MLflow Talks at Spark + AI Summit 2019 Europe

We are thrilled to see how well MLflow has been welcomed by the community since we launched it last summer. With now over 800K monthly downloads, 130 code contributors and dozens of contributing organizations including RStudio and Microsoft, it is one of the fastest growing open source projects in the field of Machine Learning, confirming the need for an open source platform to help manage the complete ML lifecycle.

We are very excited to host some of our key contributors and customers next month at the Spark + AI Summit Europe, from October 15 to 17, 2019, in Amsterdam. Below is a list of MLflow tutorials, sessions, and trainings for you to dive in.

Spark + AI Summit 2019 - Largest data & machine learning conference in the world

MLflow and Model Deployment Training

Register to Machine Learning in Production: MLflow and Model Deployment for a full-day course on MLflow, where you will learn best practices for putting machine-learning models into production.

Simplifying ML Model Management with MLflow Keynote

Join Matei Zaharia on Thursday, October 17 for a new keynote: Simplifying Model Management with MLflow to learn more about some of the most recent and new MLflow features. Specifically, we will focus on model management with the MLflow Model Registry. Many organizations face challenges tracking which models are available in the organization and which ones are in production. The MLflow Model Registry provides a centralized database to keep track of these models, share and describe new model versions, and deploy the latest version of a model through APIs.

MLflow Deep Dives and Talks

We have a fantastic lineup of speakers and sessions throughout the conference on MLflow. Join experts from TomTom, Seldon Technologies, Microsoft, Société Générale, Asurion, Databricks and more for real-life examples and deep dives on MLflow:

Free Tutorial on Managing the Complete ML Lifecycle

Last but not least, you can join Managing the Complete Machine Learning Lifecycle with MLflow for a free 80 minutes hands-on-lab presented by Thunder Shiviah and Michael Shtelma of Databricks. In this MLflow tutorial, we will show you how using MLflow can help you keep track of experiments runs and results across frameworks, quickly reproduce runs, and productionize models using Databricks production jobs, Docker containers, Azure ML, or Amazon SageMaker.

Next Steps

You can browse through our sessions from the Spark + AI 2018 Summit Europe schedule, too.

To get started with open source MLflow, follow the instructions at mlflow.org . We are excited to hear your feedback!

If you’re an existing Databricks user, you can start using Managed MLflow by importing the MLflow Quick Start Notebook for Azure Databricks or AWS. If you’re not yet a Databricks user, visit databricks.com/mlflow to learn more and start a free trial of Databricks and Managed MLflow.

Related Blogs:

 

--

Try Databricks for free. Get started today.

The post A Guide to MLflow Talks at Spark + AI Summit 2019 Europe appeared first on Databricks.

Tuesday, 24 September 2019

Audit Log Plugin for Jenkins Releases 1.0

Thanks to our Outreachy interns over the past year, I’m proud to announce the initial release of the Audit Log plugin for Jenkins. This plugin is the first major project completed related to Outreachy, and I’d like to give a brief overview of the functionality that was developed for this release. The primary goal of this plugin is to introduce an audit trail of various Jenkins events using structured logging and related audit logging standards. Initially, this plugin covers audit events related to core Jenkins concepts like user accounts, jobs, builds, nodes, and credentials usage. More specifically, this tracks:

  • User login and logout events

  • Credentials usage

  • User creation (when using the Jenkins user database as a security realm)

  • User password updates (ditto)

  • Starts and ends of builds

  • Creation/modification/deletion/copying of items (which correspond to projects, pipelines, folders, etc.)

  • Creation/modification/deletion of nodes.

This plugin defines and exports standardized log event classes and schemas corresponding to these events. Other plugins can add audit-log as a dependency to define their own audit events using Apache Log4j Audit and its catalog editor; then they can use the Maven plugin for generating the audit event classes for use in the plugin.

The other major feature of this plugin is configuring where to output these audit logs. By default, audit logs will be written in HTML files (rotated once per day) to $JENKINS_HOME/logs/html/audit.html which are viewable through the "Audit Logs" root action link. In the system settings, a section for audit logging is added where the main audit log output can be configured. This can initially be configured to output via either a JSON log file in $JENKINS_HOME/logs/audit.log by default or to a syslog server using RFC5424 encoding.

Overall, this experience has been rather interesting. Besides having an opportunity to mentor new contributors, Outreachy has helped open my eyes to the struggles that developers from around the world are dealing with which can be improved upon to help expand our communities. For example, many countries do not have reliable internet or electricity, so the use of synchronous videoconferencing and other heavyweight, synchronous processes common to more corporate-style development are inadequate in this international context. This doesn’t even begin to account for the difference in timezones which is not always an issue, though both problems are addressable by using asynchronous communication methods like chat and email. This notion of asynchronous communication is an important aspect of the Apache Way, for example, which emphasises processes that allow for vendor neutral communities to form and thrive around a project.

This mentoring project was valuable to myself as well. As a software engineer myself, project management is not my specialty, so this gave me a great opportunity to develop my own PM skills and technical leadership. My own typical discovery process for feature development involves experimenting directly with the code to see what features make sense to prioritize and which would take a vast effort to implement. Changing my own discovery process to avoid implementing the features myself was difficult to adjust to, though I did defer any of my own feature contributions to this plugin until after the initial release. In order to appropriately scope the project, I still had to spend a bit of time reading through the Jenkins codebase to determine which tasks could be implemented simply (e.g., good newbie-friendly issues), which tasks might require changes to Jenkins itself (previously discovered to take too long for these relatively short Outreachy rounds), and which tasks would require intimate familiarity with Jenkins and would likely be infeasible for new developers to Jenkins. Thanks to the work done in discovery and delivery, I’ve also identified potential features for Log4j itself which could be used in future versions of this plugin.

Overall, I think we did a good job of balancing the scope of this project without spending too much time in any specific area. The first release of this plugin is now available in the Jenkins Update Center. In the future, I hope to learn more about developing Jenkins UI components so that we can create a more dynamic and Jenkins-like configuration page for choosing where logs are output. While I don’t intend on using this plugin for further Outreachy rounds, I do hope to see more interest in it over time as the more security-conscious users out there discover this new plugin.

Backup Cams and AI Autonomous Cars

By Lance Eliot, the AI Trends Insider

[Ed. Note: For reader’s interested in Dr. Eliot’s ongoing business analyses about the advent of self-driving cars, see his online Forbes column: https://forbes.com/sites/lanceeliot/]

One of the worst nightmares for any driver is the chance of backing up a car and running over someone.

When you are backing up, it can be very difficult to know what’s behind the vehicle.

There’s a famous video of a crawling baby in Brazil that unbeknownst to the parents crawled behind the family car as it was being backed out of the garage. The car got about halfway over the top of the baby when a person walking nearby pointed out there was a baby underneath.

Getting out of the car, now stopped over the baby, the family members were fortunately able to pull out the baby and did so without much harm having come to the child.

They were lucky.

Statistics were against them in the sense that by-and-large once a backout is underway, whomever is getting hit is likely to be severely injured or even killed.

Per federal data, in the United States alone there are over 200 deaths annually and more than 15,000 persons injured via backover incidents.

As you might guess, young children in the age of 5 or less account for nearly one-third of those deaths (of those, mainly children in the 1-2 years old bracket), while adults over the age of 70 are about one-quarter of the deaths. In essence, very young children and older elders are the most likely major segments of being the victim of a backover death.

I’ve been fortunate to never backover anyone, though I’ve had a few other close calls of varying kinds.

In one case, while in college, a friend of mine thought it would be funny to hide behind my car as I was backing out, and then bang the car and act like I had hit him.

I rushed out of the car and my heart was racing as I really thought I had somehow done this. My mind was instantly thinking about my first aid training and also where the nearest hospital was. When he stood up and laughed, I assure you that I did not consider it much of a joke and was darned angry about the whole thing.

I backed over some of my children’s toys from time-to-time that they had left laying on the ground behind the car.

I wised up to this and would always check behind the car before I started to back-up. They would get trickier though and sometimes the toys were already underneath the car before I started to back-up, and once I began going back I would hear and feel the crunch of some toy getting smashed by the tires or the underbelly of the car. In each case, it made me cringe because it reminded me that it could have been perhaps a human instead of a toy.

Backup Cams: It’s The Law

I’m sure that you are thinking that if we had backup cams on cars then we wouldn’t have any of these deaths and injuries.

Maybe.

First, you might find of keen interest that after years of delays in implementing a law passed by Congress in 2008 requiring regulators to enact legal measure that would require auto makers to enhance rear view visibility, some ten years later, there is indeed a requirement that new cars being sold in the United States must be outfitted with backup cameras (the new requirement was announced in 2014 and the car makers were given four years to implement it, starting in 2018).

For those of you with modern cars, you’ve already likely got a backup cam in your car, so this law doesn’t mean much to you, other than the aspect that gradually there will be a lot of cars with the backup cam.

Eventually, once older cars end-up on the junk heap, all cars will have backup cameras as the newer cars become the dominant proportion of all 200+ million cars of today (this will take many years though to playout).

I guess we can call the “backover” problem solved since we’re going to have all these backup cams – if you believe this you are in for a bit of a surprise.

Turns out that a study in 2016 found that of cars outfitted with a backup cam and the same model of cars without a back-up cam, there was only about a 16% drop in reported backover incidents for those cars with the back-up cam.

Think about that for a moment. You might have assumed that there should be a 100% drop in backover incidents. Having a back-up cam implies no more backovers.

Backing Up Is A Driving Weakness

Well, a backup cam is only as useful as the nature of the driver at the wheel.

It is unlikely that all drivers will actually look at the display in their car to see what the backup cam shows them.

For some people, they get so used to the backup cam that they rarely look at it. I know one driver that looks over his shoulder instead of looking at the backup cam display, which he insists is a better approach than relying upon the backup cam. Good old over-the-shoulder in his book is by far superior to the “useless” backup cam.

Even if the driver does look at the backup cam, they might not notice what the backup camera display is showing them.

A baby laying on the floor behind the car might be laying still and thus there isn’t any movement shown in the display, and thus the driver doesn’t notice the child laying there. It might seem far fetched to you that someone would not notice a baby laying on the ground but imagine that you backout of your garage every morning to get to work, and some mornings you are in a rush, and 99% of the time there’s nothing at all behind your car, and you “know” that the baby is inside the house (or so you assume). All of those factors can allow someone to mindlessly not see what the display is showing.

Another factor is whether an object moves into the field of vision at the last moment.

Perhaps at first, before you start the car, you glance at the backup cam display, and mentally make a note that there’s nothing behind you. So, you put the car in reverse and you take your eyes off the display. You look over your shoulder, or maybe in the rearview mirror, and begin to backup. You are feeling confident that there’s nothing directly behind you.

Using the story of the Brazilian baby that crawled behind a car, imagine if the baby was at the sides outside the view of the cam initially, and managed to crawl behind the car at the most inopportune moment. This could happen in a few split seconds of time.

Of course, we also need to consider the field of vision of the backup cam as key to this too.

Depending upon what kind of backup cam you have, and how it is mounted, you can have a narrow view or a wide view. You can have a view that sees down to the floor, or a view that is more of an upward look. There can be blind spots that the backup cam does not show you. The backup cam can also get obscured with dirt or other obstructions. Making assumptions that the backup cam gives you an all-knowing all-seeing vantage of what’s behind the car is a mistaken belief.

Furthermore, some backup cameras allow you to have a variety of vantage viewpoints, and you as the driver need to select the one that you want to see.

Some drivers, being lazy or just not attuned to the multiple views, have a tendency to leave the display set to a particular view. The driver doesn’t rotate through them each time they do a backup operation. Most drivers become complacent about backing up when in a familiar setting. If you backup each day at your driveway or garage, you become accustomed to doing so. One sobering statistics is that by-and-large the person injured or killed is a family member or similar relation to the driver.

In essence, most of us will really only study the display of the backup cam when we get into dicey situations.

You are at the grocery store and need to backup out of your parking spot. You saw that children were wandering around the parking lot and playfully having fun. All of a sudden, you devote your attention to the backup cam. Or, you are backing up and it’s a really tight space situation, and so again you are on your alert and pay special attention to the backup cam display. The rest of the time, it’s there but you don’t put any mind toward it.

Automation To Aid The Backing Up Driver

If the backup cam alone won’t get the job done, I am sure you are thinking that let’s put some automation onto the task.

Indeed, there are some backup cam systems that have an alerting feature. If the backup cam detects an object in the field of view, it will make a tone or some other alert inside the car to let the driver know. This could definitely help for those drivers that aren’t rigorously always studying their backup cam.

In addition to an alert, or in lieu of an alert, another kind of automation is an emergency braking system for backover prevention or mitigation purposes.

If the emergency braking system detects an object within the field of view, and if the car is backing up and in motion, the emergency braking system acts like a collision avoidance system and stops the car, doing so regardless of what the driver might try to do. Automatic emergency braking systems while in forward motion are becoming increasingly common on new cars and will become an auto industry recommended requirement for all new cars and light trucks sold in the United States starting in 2022. This though does not apply to backup emergency braking systems.

Here’s the story so far.

We know that backovers are a deadly problem.

We know that having a backup cam helps to reduce the number of backovers, but doesn’t curtail it entirely.

We know that if you had an alert coupled with the backup cam, it would likely further reduce the backovers.

If there was also an emergency braking system that applied to backing up, it too would likely even further reduce backovers. As they say, we have the means, but we don’t quite yet have the willpower.

There is a cost to putting in an alert system and an emergency braking system for backover purposes, and society hasn’t reached a point where it wants those so much that it has demanded that they be put onto cars.

AI Autonomous Cars And Backing Up

What does this have to do with AI self-driving driverless autonomous cars?

At the Cybernetic Self-Driving Car Institute, we point out that by-and-large self-driving cars are going to be equipped with sensors that allow for looking behind the car, and so it is an already built-in capability that just needs to be leveraged by the AI of the self-driving car. That’s an “edge” problem that we are working on.

An edge problem is considered one that is not at the core of something.

The core of an AI self-driving car is having the AI be able to drive the car forwards, and be able to drive down streets, drive on freeways, make safe left turns, and otherwise do all the things a human driver can do. The true self-driving car is considered a Level 5, meaning that it is a self-driving car that is driven entirely by the AI without any human intervention, and that the AI can drive as a human could.

See my article on the levels of self-driving cars: https://aitrends.com/selfdrivingcars/richter-scale-levels-self-driving-cars/

In essence, right now, the auto makers and tech firms are mainly concerned with making an AI self-driving car that can drive forwards.  Going in reverse is considered a secondary problem, or what some refer to as an edge problem, since its not at the core of the driving task (as they view it). Yes, it is an important part of driving, but it’s not as crucial as driving forwards, they contend. In fact, some of the AI developers consider going in reverse to be fully solvable by just taking what you’ve developed for going forwards and reapplying it when the car is in reverse.

We don’t believe that going in reverse is merely the same as going forwards, see my article on this topic: https://aitrends.com/selfdrivingcars/looking-behind-self-driving-car-neglected-blind-spot/

This also brings up the important point that if you already are wanting AI self-driving cars so as to reduce the estimated 40,000 annual deaths in the United States due to driving incidents by human drivers, you would presumably also like to see that 15,000 per year injuries due to backing up would also get reduced too.

The AI self-driving car pretty much should already have the needed sensory devices, and so the other aspect is the AI software to leverage those sensors.

That being said, not all of the emerging AI self-driving cars necessarily have a typical backup cam per se.

They might have other cameras on the rear of the vehicle, though not necessarily of the type for a backup purpose and nor aimed at the ground behind the vehicle.

Some of the cameras are instead aimed at a further distance, so as to detect a car behind the self-driving car. One added potential plus for an AI self-driving car is that there are usually radar, sonar, and LIDAR on the car too, which can be used in combination with the cameras.

I want to point out that very important element that I just mentioned.

A conventional car that is outfitted with a backup cam is unlikely to have radar, sonar, LIDAR, and other sensors that can be used in combination with the backup cam. A conventional backup cam is all alone. It is the only means to try and detect what’s behind the car. This is slim. Having only visual clues about what is behind a car can be misleading or distorted. It is advantageous to have multiple ways to detect what’s behind the car.

We’re incorporating the other sensors into the gambit of preventing backovers so that we can increase the chances of avoiding a backover, doing so by bringing together the visual data, the radar data, the sonar data, the LIDAR data, and the rest. The AI of the self-driving car has to be doing some solid defensive driving when backing up.

For more info about AI self-driving car defensive driving, see my article: https://aitrends.com/selfdrivingcars/art-defensive-driving-key-self-driving-car-success/

Why Backing Up In Cars Is Vital

It is helpful to consider the major use cases associated with backing up a car.

We back out a car in relatively common circumstances.

There are exceptions beyond the common circumstances, but it’s best to focus initially on the common ones and then branch out from there.

First, there is backing out of a garage as a driving-in-reverse type of task.

This is extremely common for people to want to do.

The task is not as easy as it might seem. Sometimes a garage has a lot of junk in it and the sensors cannot detect anything distinctive (it’s one large blur). A garage can have very tight quarters and it makes the sensors unable to work appropriately. Garage parking and getting in and out is a specialized kind of problem (another edge problem!).

See my article about garage parking: https://aitrends.com/selfdrivingcars/ai-home-garage-automatic-parking-self-driving-cars/

Next, there’s the act of driving down a driveway while backing out.

A similar use case is driving up a driveway while backing out.

There’s backing into a parking spot as another commonly performed task, and likewise backing out of a parking spot (less frequent, but sometimes combined with going back-and-forth to inch out of a tight parking spot).

For the particulars about AI self-driving cars and parallel parking, see my article: https://aitrends.com/selfdrivingcars/parallel-parking-mindless-ai-task-self-driving-cars-time-step/

Throughout any of those backing up operations, the AI needs to be on the watch for obstructions.

The obstructions might be large or small. They might be still or in motion. They might be readily detected by multiple sensors, or only detected by one of the sensors. They might be moving away from the self-driving car or toward the self-driving car. It could be an obstruction that is recognizable, such as perhaps detecting that the obstruction is a person, or it might be an obstruction that is unrecognizable, which is nonetheless an obstruction but one of unknown capabilities or purposes.

There are also some important exceptions.

For example, suppose there is a person purposely standing at the back of my AI self-driving car that is warning other people to stay clear.

This helpful person, they themselves now become a detected obstruction by the AI. How will the AI know that the obstruction is actually part of the backing up operation?

Instead, the AI is going to assume that the person standing there is to be avoided and will likely bring the car to a halt.

Speaking of which, every morning there is a newspaper tossed onto my driveway.

Exceptions To Backing Up That Allow Driving Over Something

Each morning, I dutifully back down the driveway and go over the newspaper.

When I get home at night, I once again drive over the newspaper, and upon parking my car in the garage, I get out of my car and go to get the newspaper on the driveway.

Let’s for the moment assume that the self-driving car sensors and the AI are good enough to be able to detect that the newspaper is sitting there on the driveway.

The AI would  presumably refuse each morning to backup and later when I get home would refuse to go forward, since in both cases I am driving over something.

Thus, this is a harder problem than it might seem, since there are going to be circumstances where the human occupants want the AI self-driving car to proceed with backing up, in spite of a potential rollover of something or the nearness of a human or other object.

The human occupant will need to have some means to communicate with the AI self-driving car, such as by using an in-car command. This human directive capability though has a downside, since suppose the human occupant intentionally wants to harm someone and so tells the AI self-driving car to proceed to backup into the person – should the AI self-driving car comply? As you can see, there are ethics issues involved in this too.

See my article about ethics and AI self-driving cars: https://aitrends.com/selfdrivingcars/ethically-ambiguous-self-driving-cars/

Conclusion

We know that backup cams are coming to conventional cars, slowly, gradually, and that some cars will also have alerts or emergency braking systems combined with a backup cam.

Older cars are unlikely to have this.

Only some of the newer cars will have it.

For AI self-driving cars, they are destined from the start to have sensory devices that can be leveraged for backing up safely.

We just need to make sure that right kind of sensors are being included, and that the AI is savvy enough to leverage those sensors.

I’d like to be able to say that with AI self-driving cars we’ll have eliminated the backover problem, but realistically it won’t eliminate it, but at least it should help to reduce the frequency and magnitude of backover incidents.

Let’s not back out of that goal.

Copyright 2019 Dr. Lance Eliot

This content is originally posted on AI Trends.

Diving Into Delta Lake: Schema Enforcement & Evolution

Try this notebook series in Databricks

Think back to when you were in high school – so fresh and full of ideas. Since then, undoubtedly, both the world and the way you see things have changed in many ways, as you have gained new experiences.

Data, like our experiences, is always evolving and accumulating. To keep up, our mental models of the world must adapt to new data, some of which contains new dimensions – new ways of seeing things we had no conception of before. These mental models are not unlike a table’s schema, defining how we categorize and process new information.

This brings us to schema management. As business problems and requirements evolve over time, so too does the structure of your data. With Delta Lake, as your data changes, incorporating new dimensions is easy. Users have access to simple semantics to control the schema of their tables. These tools include schema enforcement, which prevents users from accidentally polluting their tables with mistakes or garbage data, as well as schema evolution, which enables them to automatically add new columns of rich data when those columns belong. In this blog, we’ll dive into the use of these tools.

Understanding Table Schemas

Every DataFrame in Apache Spark™  contains a schema, a blueprint that defines the shape of the data, such as data types and columns, and metadata. With Delta Lake, the table’s schema is saved in JSON format inside the transaction log.

What Is Schema Enforcement?

Schema enforcement, also known as schema validation, is a safeguard in Delta Lake that ensures data quality by rejecting writes to a table that do not match the table’s schema. Like the front desk manager at a busy restaurant that only accepts reservations, it checks to see whether each column in data inserted into the table is on its list of expected columns (in other words, whether each one has a “reservation”), and rejects any writes with columns that aren’t on the list.

How Does Schema Enforcement Work?

Delta Lake uses schema validation on write, which means that all new writes to a table are checked for compatibility with the target table’s schema at write time. If the schema is not compatible, Delta Lake cancels the transaction altogether (no data is written), and raises an exception to let the user know about the mismatch.

To determine whether a write to a table is compatible, Delta Lake uses the following rules. The DataFrame to be written:

  • Can not contain any additional columns that are not present in the target table’s schema. Conversely, it’s OK if the incoming data doesn’t contain every column in the table – those columns will simply be assigned null values.
  • Cannot have column data types that differ from the column data types in the target table. If a target table’s column contains StringType data, but the corresponding column in the DataFrame contains IntegerType data, schema enforcement will raise an exception and prevent the write operation from taking place.
  • Can not contain column names that differ only by case. This means that you cannot have columns such as ‘Foo’ and  ‘foo’ defined in the same table. While Spark can be used in case sensitive or insensitive (default) mode, Delta is case-preserving but insensitive when storing the schema. Parquet is case sensitive when storing and returning column information. To avoid potential mistakes, data corruption or loss issues (which we’ve personally experienced at Databricks), we decided to add this restriction.

To illustrate, take a look at what happens in the code below when we attempt to append some newly calculated columns to a Delta Lake table that isn’t yet set up to accept them.


# Generate a DataFrame of loans that we'll append to our Delta Lake table
loans = sql("""
            SELECT addr_state, CAST(rand(10)*count as bigint) AS count,
            CAST(rand(10) * 10000 * count AS double) AS amount
            FROM loan_by_state_delta
            """)

# Show original DataFrame's schema
original_loans.printSchema()
 
"""
root
  |-- addr_state: string (nullable = true)
  |-- count: integer (nullable = true)
"""
 
# Show new DataFrame's schema
loans.printSchema()
 
"""
root
  |-- addr_state: string (nullable = true)
  |-- count: integer (nullable = true)
  |-- amount: double (nullable = true) # new column
"""
 
# Attempt to append new DataFrame (with new column) to existing table
loans.write.format("delta") \
           .mode("append") \
           .save(DELTALAKE_PATH)

""" Returns:

A schema mismatch detected when writing to the Delta table.
 
To enable schema migration, please set:
'.option("mergeSchema", "true")\'
 
Table schema:
root
-- addr_state: string (nullable = true)
-- count: long (nullable = true)
 
 
Data schema:
root
-- addr_state: string (nullable = true)
-- count: long (nullable = true)
-- amount: double (nullable = true)
 
If Table ACLs are enabled, these options will be ignored. Please use the ALTER TABLE command for changing the schema.

"""

Rather than automatically adding the new columns, Delta Lake enforces the schema and stops the write from occurring. To help identify which column(s) caused the mismatch, Spark prints out both schemas in the stack trace for comparison.

How Is Schema Enforcement Useful?

Because it’s such a stringent check, schema enforcement is an excellent tool to use as a gatekeeper of a clean, fully transformed data set that is ready for production or consumption. It’s typically enforced on tables that directly feed:

  • Machine learning algorithms
  • BI dashboards
  • Data analytics and visualization tools
  • Any production system requiring highly structured, strongly typed, semantic schemas

In order to prepare their data for this final hurdle, many users employ a simple “multi-hop” architecture that progressively adds structure to their tables. To learn more, take a look at the post entitled Productionizing Machine Learning With Delta Lake.

Of course, schema enforcement can be used anywhere in your pipeline, but be aware that it can be a bit frustrating to have your streaming write to a table fail because you forgot that you added a single column to the incoming data, for example.

Preventing Data Dilution

At this point, you might be asking yourself, what’s all the fuss about? After all, sometimes an unexpected “schema mismatch” error can trip you up in your workflow, especially if you’re new to Delta Lake. Why not just let the schema change however it needs to so that I can write my DataFrame no matter what?

As the old saying goes, “an ounce of prevention is worth a pound of cure.” At some point, if you don’t enforce your schema, issues with data type compatibility will rear their ugly heads – seemingly homogenous sources of raw data can contain edge cases, corrupted columns, misformed mappings, or other scary things that go bump in the night. A much better approach is to stop these enemies at the gates – using schema enforcement – and deal with them in the daylight rather than later on, when they’ll be lurking in the shadowy recesses of your production code.

Schema enforcement provides peace of mind that your table’s schema will not change unless you make the affirmative choice to change it. It prevents data “dilution,” which can occur when new columns are appended so frequently that formerly rich, concise tables lose their meaning and usefulness due to the data deluge. By encouraging you to be intentional, set high standards, and expect high quality, schema enforcement is doing exactly what it was designed to do – keeping you honest, and your tables clean.

If, upon further review, you decide that you really did mean to add that new column, it’s an easy, one line fix, as discussed below. The solution is schema evolution!

What Is Schema Evolution?

Schema evolution is a feature that allows users to easily change a table’s current schema to accommodate data that is changing over time. Most commonly, it’s used when performing an append or overwrite operation, to automatically adapt the schema to include one or more new columns.

How Does Schema Evolution Work?

Following up on the example from the previous section, we can easily use schema evolution to add the new columns that were previously rejected due to a schema mismatch. Schema evolution is activated by adding  .option('mergeSchema', 'true') to your .write or .writeStream Spark command.


# Add the mergeSchema option
loans.write.format("delta") \
           .option("mergeSchema", "true") \
           .mode("append") \
           .save(DELTALAKE_SILVER_PATH)

To view the plot, execute the following Spark SQL statement.


# Create a plot with the new column to confirm the write was successful
%sql
SELECT addr_state, sum(`amount`) AS amount
FROM loan_by_state_delta
GROUP BY addr_state
ORDER BY sum(`amount`)
DESC LIMIT 10

Alternatively, you can set this option for the entire Spark session by adding spark.databricks.delta.schema.autoMerge = True to your Spark configuration. Use with caution, as schema enforcement will no longer warn you about unintended schema mismatches.

By including the mergeSchema option in your query, any columns that are present in the DataFrame but not in the target table are automatically added on to the end of the schema as part of a write transaction. Nested fields can also be added, and these fields will get added to the end of their respective struct columns as well.

Data engineers and scientists can use this option to add new columns (perhaps a newly tracked metric, or a column of this month’s sales figures) to their existing machine learning production tables without breaking existing models that rely on the old columns.

The following types of schema changes are eligible for schema evolution during table appends or overwrites:

  • Adding new columns (this is the most common scenario)
  • Changing of data types from NullType -> any other type, or upcasts from ByteType -> ShortType -> IntegerType

Other changes, which are not eligible for schema evolution, require that the schema and data are overwritten by adding .option("mergeSchema", "true"). Those changes include:

  • Dropping a column
  • Changing an existing column’s data type (in place)
  • Renaming column names that differ only by case (e.g. “Foo” and “foo”)

Finally, with the upcoming release of Spark 3.0, explicit DDL (using ALTER TABLE) will be fully supported, allowing users to perform the following actions on table schemas:

  • Adding  columns
  • Changing column comments
  • Setting table properties that define the behavior of the table, such as setting the retention duration of the transaction log

How is Schema Evolution Useful?

Schema evolution can be used anytime you intend   to change the schema of your table (as opposed to where you accidentally added columns to your DataFrame that shouldn’t be there). It’s the easiest way to migrate your schema because it automatically adds the correct column names and data types, without having to declare them explicitly.

Summary

Schema enforcement rejects any new columns or other schema changes that aren’t compatible with your table. By setting and upholding these high standards, analysts and engineers can trust that their data has the highest levels of integrity, and reason about it with clarity, allowing them to make better business decisions.

On the flip side of the coin, schema evolution complements enforcement by making it easy for intended schema changes to take place automatically. After all, it shouldn’t be hard to add a column.

Schema enforcement is the yin to schema evolution’s yang. When used together, these features make it easier than ever to block out the noise, and tune in to the signal.

 

We’d also like to thank Mukul Murthy and Pranav Anand for their contributions to this blog.

Related Articles

Productionizing Machine Learning With Delta Lake

Other articles in this series:

Diving Into Delta Lake: Unpacking the Transaction Log

--

Try Databricks for free. Get started today.

The post Diving Into Delta Lake: Schema Enforcement & Evolution appeared first on Databricks.

What is the Best Architecture for your Application: Monolith, Microservices or Serverless?

Choosing the right architecture is critical for the overall success of your product. The three most popular architectures used in the IT world are Monolith, Microservices, and Serverless. Each one offers its own advantages to create exactly the right sort of solution for your users with the best possible experience. Let’s take a look at each architecture separately to see how they work and uncover potential benefits. 

Monolithic Architecture

Monolithic software is self-contained and all of the various components are interconnected with each other. This means that each component and all of the components associated with it need to be present for the code to be executed and compiled. If you would like to update even one of the components, this means that you will need to rewrite the entire application. While this scares away a lot of developers from creating monolithic architectures, it offers substantial benefits as well such as:


Fewer issues affecting the entire app -  These include error handling, logging and caching. All of these functionalities would only concern a single app thus making life simpler. 
Simpler testing and debugging - the monolith is a big cohesive unit, thus making it possible to perform end-to-end testing faster. 
Easy to deploy - Since ...


Read More on Datafloq

The Need for Whistleblowers in New Tech

Whistleblowers are the people who notice wrongdoing and put their jobs and reputations on the line by seeking to expose those at fault. They are valuable in any industry, but are particularly applicable to the tech sector.

The Tech Sector Is Fast-Moving

Succeeding when developing a tech product typically means being the first to release it to the marketplace. Sometimes, that means companies conceal certain things. Representatives at those firms know that to reveal them would hinder the ability to get funding from investors, and it would likely tarnish the public's view of the business.

The now-infamous tale of Theranos, the medical company that claimed to perform a range of tests with just a drop of a person's blood, illustrates what can happen. Erika Cheung was among those who came forward to reveal how the machine never worked as intended. Cheung said that before she began her job, the founder, Elizabeth Holmes, assured her she'd understand everything about her role after starting work.

However, Cheung soon discovered that Holmes kept employees, board members, investors and others in the dark about what really happened at the company. Even so, Cheung gathered enough information to show that the promises were too good to be true. Throughout the business's ...


Read More on Datafloq

4 Ways You're Sharing Too Much Information and How To Prevent Them

For years now, your personal data has been under attack by a variety of sources, and the problem is only getting worse. But while most people relate this issue to cybercrime activity and direct identity targeting, the reality is, many of us freely divulge too much information about ourselves in our everyday lives without even realizing it. 

Society now thrives off of the efficiency of digital data storage and accessibility. Smart devices, big data analytics systems, and AI-enabled technologies continue to provide many improvements in how we live our lives and play a major role in advancing personal and business data security.

However, it’s important to understand the dangers of allowing too much access to your personal information and how you use certain services and applications. Here are four areas where you could reveal too much information about yourself and what you can do to keep your data protected.

Social Media Profiles

Social media platforms like Facebook, Instagram, and Twitter have become a major part of people's lives and are now a preferred method of communication with friends and family. But while social media profiles give us the opportunity to express our likes, interests, and other important traits of our personalities, they are also ...


Read More on Datafloq

Monday, 23 September 2019

How to Build Data Culture and Make Data Your Friend

Sometimes working with big data resembles dining at an “all-you-can-eat” restaurant: you get too much food and are busy swallowing it all without thinking of its quality or your health. You should never forget, though, that data is a double-edged sword and requires cautious handling. If gives businesses new power to organize, operate, predict and create value, but it entails numerous risks, including data-security questions, privacy concerns, and uncertainty about ethical boundaries, to name just a few.

To create a data-driven culture is becoming critical in times of global connectivity and data-driven organizations.

“Data culture” is a relatively new concept which is becoming pivotal nowadays, when organizations develop more progressive digital business strategies and apply meaning to big data. It refers to a workplace environment that employs a consistent approach to decision-making through emphatic and empirical data proof. In other words, it implies that decisions are made based on data evidence, not on gut instinct.

Key Concepts of Data Culture



1. Data culture is a decision-making tool

Data analysis is not invented for the love of science; its fundamental objective is collecting, processing and deploying data to make better decisions with a focus on the outcomes. The insights, ideas, and innovation generated by the team shall be ...


Read More on Datafloq

4 Business Applications of Natural Language Processing

The future of tomorrow belongs to NLP. The business market has grown to such an extent where it now needs to analyze and understand customer behaviour, their preferences, as well as their mood. Had it not been for natural language processing businesses owners were likely to remain incompetent in handling and extracting valuable insights from texts.

Due to the rise of machine learning and artificial intelligence, the advent of NLP came into existence. Gartner claims conversational analytics as the newly born paradigm. There is no denying fact stating that data-driven is going to be the new mantra in the business domain. The NLP is also said to be recognized to be the enabler of text analysis as well as speech recognition applications. NLP experts will be in demand in the coming future. For professionals looking to enter the domain, it is a seller market.

Natural Language Processing is experiencing a whole new level of transformation in the field of business and service providers. Let us check out these top applications of NLP technology you should consider for your business today.

Chatbots

Perhaps chatbots are the most common application we’ve frequently come across. Chatbots are solutions for customers’ frustration over providing services. Day-to-day assistance service ...


Read More on Datafloq

How Artificial Intelligence Is Revolutionizing Insurance

Artificial intelligence has been named a disruptive force in multiple areas, including finance, healthcare, and security. The insurance sector can benefit significantly from these advancements of cognitive technology too. This is made possible with the heaps of data collected by insurance companies and not used to their full potential.

In the insurance business vertical, AI can have a positive impact at every level, from automating call center request processing to helping make accurate assessments and executive-level decisions.

Through its power to recognize patterns and anticipate actions, AI can provide a predictive environment where risks are anticipated and hedged.

Applications of AI for Insurance

There are numerous ways to use AI in the insurance industry. So far, it seems that the main areas of AI application in insurance include customer experience (58%), process optimization (43%) and product innovation (19%), from a 2018 study by Everest Global.

Ilya Kirillov, CEO of InData Labs, comments:

Expert consulting is an essential part of every successful AI project. At this stage, professionals can help alleviate difficulties in drawing up a custom AI solution development plan and outline business-focused functionalities.

Fraud Prevention

A report by the FBI shows that the cost of insurance fraud is estimated at more than $40 billion per year. The ...


Read More on Datafloq

Sunday, 22 September 2019

Friday, 20 September 2019

The Database of Tomorrow: The Self-Driving, Autonomous Database

This article is sponsored by Oracle - redefining data management with the world’s first autonomous database. 

In the coming years, the amount of data we create worldwide will grow to 175 zettabytes of data per year by 2025, up from 33 zettabytes in 2018. Over half of this data will be created by the Internet of Things devices and over 60% of it will be enterprise data. By 2025, 30% of all the data created will be in real-time, offering organisations great opportunities to constantly optimise their business.

Clearly, the organisation of tomorrow is a data organisation. However, simply collecting vast amounts of data is not enough. You would also need to analyse the data for insights and change your organisational culture to benefit from it. According to McKinsey, data-driven organisations are 23x more likely to acquire customers, 6x more likely to retain customers and 19x more likely to be profitable. Being data-driven is good for business.

The Importance of Data Governance

When collecting petabytes of data, it becomes vital that this data is of high-quality. Organisations that focus on high-quality data are better able to deal with changing business environments and achieve strategic objectives. As such, in today’s data-driven world, data governance has ...


Read More on Datafloq

The Database of Tomorrow: The Self-Driving, Autonomous Database

This article is sponsored by Oracle - redefining data management with the world’s first autonomous database. 

In the coming years, the amount of data we create worldwide will grow to 175 zettabytes of data per year by 2025, up from 33 zettabytes in 2018. Over half of this data will be created by the Internet of Things devices and over 60% of it will be enterprise data. By 2025, 30% of all the data created will be in real-time, offering organisations great opportunities to constantly optimise their business.

Clearly, the organisation of tomorrow is a data organisation. However, simply collecting vast amounts of data is not enough. You would also need to analyse the data for insights and change your organisational culture to benefit from it. According to McKinsey, data-driven organisations are 23x more likely to acquire customers, 6x more likely to retain customers and 19x more likely to be profitable. Being data-driven is good for business.

The Importance of Data Governance

When collecting petabytes of data, it becomes vital that this data is of high-quality. Organisations that focus on high-quality data are better able to deal with changing business environments and achieve strategic objectives. As such, in today’s data-driven world, data governance has ...


Read More on Datafloq

Why Data Privacy is a Crucial Part of Customer Experience?

The world runs on data. And businesses that collect and process them are extremely valuable and powerful. But then, no one has the right to intrude on data privacy and violate their customers' trust —  no, not even the massive brands that drive the global economy.

Notorious incidents like the Wells Fargo scandal in 2016 led us to believe that we live in the worst time of data security. 

Here’s what happened. Employees at Wells Fargo were accused of opening more than 2 million unauthorized bank and credit card accounts without informing their customers about it. The now-former CEO John Stumpf came under strict scrutiny for fostering such high-pressure cross-selling practices that spawned the scam. 

The world has dramatically changed over the last 24 years. Misuse of internet and personal data today is fundamentally different than it was in the past. Therefore, it became crucial to replace the outdated Data Protection Directive (enacted back in 1995) with something more stringent and rigid for the 21st century. 

What is GDPR and what does it mean for your enterprise? 

GDPR or General Data Protection Regulation is a set of rules regularized by the European Union in 2018 to define the digital privacy of its citizens. At its ...


Read More on Datafloq

Where Is Cybersecurity Headed for Autonomous Vehicles?

Giving up control of our automobiles to a potentially fallible computer is a foreign concept to many of us. And yet, that’s what we do each time we climb behind the wheel of a car or truck. The trouble is, the computer between our ears is a lot slower and less consistent than the ones we build from silicon.

Autonomous cars will take some getting used to, even as they help us bring down the number of crashes and fatalities on our roads. But cybersecurity is a big part of the learning curve here — maybe even more so than convincing people to give up the wheel. Here’s a look at where cybersecurity for autonomous vehicles is headed.

When Autonomous Cars Become Tools for Cyber Criminals

Researchers representing the U.S. Department of Transportation estimate that fully driverless cars could reduce fatalities on the road by as much as 94%. This is an unprecedented opportunity to save millions of lives over the coming years.

Driverless cars rely on software and hardware to engage in safe pathfinding. As more of these cars make their way onto our roads, we greatly increase our cybersecurity threat surface. Like any mobile computer, autonomous cars need to exchange data with ...


Read More on Datafloq

Engineering population scale Genome-Wide Association Studies with Apache Spark, Delta Lake, and MLflow

Try this notebook series in Databricks

The advent of genome-wide association studies (GWAS) in the late 2000s enabled scientists to begin to understand the causes of complex diseases such as diabetes and Crohn’s disease at their most fundamental level. However, academic bioinformatics tools to perform GWAS have not kept pace with the growth of genomic data, which has been doubling globally every seven months.

Given the scale of the challenge and the importance of genomics to the future of healthcare, at Databricks we have dedicated an engineering team to develop extensible Spark-native implementations of workflows such as GWAS, which leverage the high performance big-data store, Delta Lake, and log runs with MLflow. Combining these three technologies with a library we have developed in-house to enable customers to work with genomic data solves the challenges that we have seen our customers face when working with population-scale genomic data.

This tooling includes an architecture that allows users to ingest genomics data directly from flat file formats such as bed, VCF, or BGEN, into Delta Lake. In this blog, we focus on moving common association testing kernels into Spark SQL, streamlining the running of common tests such as genome-wide linear regression.  In our next blog, we will generalize this process by using the pipe-transformer parallelize any single-node bioinformatics tool with Spark, starting with the GWAS tool SAIGE.

Here we showcase how to run and end-to-end GWAS workflow in a single notebook using the publicly available 1,000 genomes dataset, producing the results in figure 1. We used associated variants from the GWAS catalog to generate a synthetic body-mass index (BMI) phenotype (since the 1000 Genomes project did not capture phenotypes). This notebook is written in Python, but can also be implemented in R, Scala and SQL.

Figure 1. Databricks dashboard showing key results from a GWAS on simulated data based on the 1000 genomes dataset.

Ingest 1,000 Genomes Data into Delta Lake

To start, we will load in the 1,000 Genomes VCF file as a Spark SQL DataFrame and calculate summary statistics. Our schema is an intuitive representation of genomic variants that is consistent across both VCF and BGEN data.


# Reading Databricks Delta version of the VCF file
spark.read.format("com.databricks.vcf"). \
           option("splitToBiallelic", "true"). \
           option("flattenInfoFields", "false"). \
           load(vcf_path). \
           selectExpr("*", "expand_struct(call_summary_stats(genotypes))", "expand_struct(hardy_weinberg(genotypes))"). \
           write. \
           format("delta"). \
           save(delta_path)

Figure 2. Databricks’ display() command showing VCF file in a Spark DataFrame

The 1,000 Genomes dataset contains whole genome sequencing data, and thus includes many rare variants. By running a count query on the dataset, we find that there are more than 80 million variants. Let’s go ahead and log this metric to MLflow.


# VCF count
num_variants = spark.read.format("delta").load(delta_path).count()
mlflow.log_metric("Number Variants pre-QC", num_variants)
num_variants

# Output
81271745

Perform quality control

In our genomics library, we have added quality control functions that compute common statistics across the genotypes at a single variant, as well as across all of the samples in a single callset. Here we are going to filter variants that are not in Hardy-Weinberg equilibrium (“pValueHwe”), which is a population genetics statistic that can be used to assess if variants have been correctly genotyped. We will exclude rare variants based on allele frequency.


spark.read.format("delta"). \
   load(delta_path). \
   where((col("alleleFrequencies").getItem(0) >= allele_freq_cutoff) & 
         (col("alleleFrequencies").getItem(0) <= (1.0 - allele_freq_cutoff)) & (col("pValueHwe") >= hwe_cutoff)). \
   write. \
   format("delta"). \
   save(delta_qc_path)

Figure 3. Histogram of Hardy-Weinberg Equilibrium P values

Control for ancestry

Population structure can confound genotype-phenotype association analyses. To control for differing ancestry between participants in the study, here we calculate principal components (PCs), which are provided as covariates to the regression kernel. Spark supports singular value decomposition (SVD) through the Spark MLLib DistributedMatrix API, and SVD can be used to calculate PCs from the transpose of the genotypes matrix. We have introduced an API in Spark that makes it easy to build a DistributedMatrix from a DataFrame, and use this to run SVD and get our PCs.


vectorized = spark.read.format("delta"). \
                        load(delta_qc_path). \
                        selectExpr("array_to_sparse_vector(genotype_states(genotypes)) as features").cache()

matrix = RowMatrix(MLUtils.convertVectorColumnsFromML(vectorized, "features").rdd.map(lambda x: x.features))
pcs = matrix.computeSVD(num_pcs)
pcs_df = spark.createDataFrame(pcs.V.toArray().tolist(), ["pc" + str(i) for i in range(num_pcs)])

After running PCA, we get back a dense matrix of PCs per sample, that we will pass as covariates to the regression analysis. The next steps will extract out only the sampleId and the principal components. This allows us to join against the 1,000 Genomes sample metadata file to label each sample with their super-population.

With Databricks’ display() command, we can view the clusters of our components within the following scatterplot.

Figure 4. Principal Component Analysis with Super Population Labelling:

EUR = European, EAS = East Asian, AMR = Admixed American, SAS = South Asian, AFR = African

Ingest Phenotype Data

For this genome-wide association study, we will be using simulated BMI phenotypic data to associate with the genotypes. Similar to the ingestion of our genotype data, we will ingest the BMI data by reading our sample Parquet data.


# Ingest normalized phenotype data
phenotypes_path = "dbfs:/databricks-datasets/genomics/1000G/phenotypes.normalized"
bmiPhenotype = spark.read. \
                     format("parquet"). \
                     load(phenotype_path). \
                     withColumnRenamed("values", "phenotype_values")

# View BMI data
display(bmiPhenotype.selectExpr("explode(phenotype_values) AS bmi"))

You can visualize the BMI histogram from the preceding display() command.

Figure 5. BMI histogram

Running the Genome-Wide Association Study

Now we have performed the necessary quality control and data extraction, transformation and loading (ETL), the next phase of our solution is to run our GWAS by performing the following tasks:

  • Mapping the genotypes, phenotypes, and principal components together (using crossJoin).
  • Calculate the GWAS statistics by running linear regression.
  • Build a new Apache Spark DataFrame (gwas_df) that contains the GWAS statistics.

# Map variants to GWAS via cross-joins between genotypes, phenotypes, and principle components
covariates = spark.read.format("delta").load(principal_components_path)
phenotypeAndCovariates = bmiPhenotype.crossJoin(covariates)
genotypes = spark.read.format("delta").load(delta_qc_path)

genotypes.crossJoin(phenotypeAndCovariates). \
          selectExpr("contigName", "start", "phenotype", \
                     "expand_struct(linear_regression_gwas(genotype_states(genotypes), phenotype_values, covariates))"). \
          write. \
          format("delta"). \
          save(gwas_results_path)

# Display data
display(spark.read.format("delta").load(gwas_results_path))

Figure 6. Spark DataFrame of GWAS results

The display() command allows us to sanity check the results. Next we can convert our PySpark DataFrame to R thus allowing us to use the qqman package to visualize the results across the genome with a Manhattan plot.


# Extract out GWAS results (and alias various column names)
gwas_results <- select(gwas_df, c(cast(alias(gwas_df$contigName, "CHR"), "double"), alias(gwas_df$start, "BP"), "P"))

# Convert from a Spark DataFrame to an R DataFrame
gwas_results_rdf <- as.data.frame(gwas_results)

# Install packages necessary for Manhattan plot
install.packages("qqman", repos="http://cran.us.r-project.org")
library(qqman)

# Create Manhatatan plot of GWAS results and log to MLflow
png('/databricks/driver/manhattan.png')
manhattan(gwas_results_rdf, 
          col = c("#228b22", "#6441A5"), 
          chrlabs = NULL,
          suggestiveline = -log10(1e-05), 
          genomewideline = -log10(5e-08),
          highlight = NULL, 
          logp = TRUE, 
          annotatePval = NULL, 
          ylim=c(0,17))
dev.off()

mlflow.log_artifact('/databricks/driver/manhattan.png')

Figure 7. GWAS Manhattan Plot

As you can see from our genome-wide association study, for our 1000 genomes simulated data, there are several loci associated with BMI clustered on chromosome 2. In fact, these are the loci whose known associations with BMI were used to simulate our BMI phenotype.


# Execute QQ plot
qq(gwas_results_rdf$P)

We can also check that we have successfully controlled for ancestry by making a quantile-quantile (QQ) plot. In this case, the deviation from expected represents true associations.

Figure 8. GWAS QQ Plot

Finally, we have logged parameters, metrics and plots associated with this GWAS run using MLflow, enabling tracking, monitoring and reproducing of analyses.

Summarizing the Analysis

In this blog, we have demonstrated an end-to-end GWAS workflow using Apache Spark, Delta Lake, and MLflow. Whether you are validating the accuracy of a genotyping assay in clinical use like Sanford Health, or performing a meta-analysis of GWAS results for target identification like Regeneron, the Databricks platform makes it easy to extend analyses and build downstream exploratory visualizations through our built-in dashboarding functionality or through our optimized connectors to BI tools like Tableau and PowerBI which can enable non-coding bench scientists and clinicians to rapidly explore large datasets.

Figure 9. MLflow tracking of each run enables reproducibility of experiments

By robustly engineering an end-to-end GWAS workflow, scientists can move away from ad hoc analysis on flat files, to scalable and reproducible computational frameworks in production. Furthermore by reading VCF data as a Spark DataSource into Delta Lake, data scientists can now integrate tabular phenotypes, Electronic Health Record (EHR) extracts, images, real-world evidence and lab values under a unified framework. Try it yourself today by downloading the Engineering population scale GWAS Databricks notebook.

Try it!

Run our scalable GWAS workflow on the Databricks platform (Azure | AWS). Learn more about our genomics solutions in the Databricks Unified Analytics for Genomics and try out a preview today.

--

Try Databricks for free. Get started today.

The post Engineering population scale Genome-Wide Association Studies with Apache Spark, Delta Lake, and MLflow appeared first on Databricks.

ISRO panel to look into Vikram loss

Space agency’s attempts to communicate with the lander have been futile till date

Electrifying opportunity, some bumps on the way

A slew of researchers and entrepreneurs are working overtime to reduce battery costs and improve efficiency, while yet others are looking to put new models of electric vehicles on the road.

White House Sets $973M Non-Defense AI R&D Budget for FY2020

The White House recently released its FY2020 non-defense AI R&D spending request, totaling just under $1 billion.

Michael Kratsios, U.S. CTO and head of the White House’s Office of Science and Technology Policy, announced the report during a recent speech at the Information Technology and Innovation Foundation’s Center for Data Innovation event. He said the budget request is practical. “In American AI R&D budgets, you won’t find aspirational expenditures or cryptic funding mechanisms,” he said in an account in MeriTalk. “Our future rests on getting AI right,” he added.

He mentioned how authoritarian governments are using AI among technologies used to control their people, partly by limiting free speech. “This is not the American way,” he said.

See the source article in MeriTalk.

Road Diets and AI Autonomous Cars

By Lance Eliot, the AI Trends Insider

[Ed. Note: For reader’s interested in Dr. Eliot’s ongoing business analyses about the advent of self-driving cars, see his online Forbes column: https://forbes.com/sites/lanceeliot/]

There must be hundreds or maybe even thousands of different diet regimens that you can opt to use.

Some diets might be good for you, while others can have adverse consequences that outweigh the benefits of the dietary aspects.

People often struggle to choose the right diet for them, and they equally tend to struggle trying to stay on the diet and stick with it. I’ve known some people that became quite irritable and onerous once they got onto a diet program.

Let’s consider another kind of diet, namely a road diet.

If you aren’t familiar with the notion of a road diet, don’t feel too bad about not knowing what it is.

In some sense, road diets are a bit of a fad that seems to come and go.

Generally, a road diet consists of taking an existing roadway and altering it to reduce the number of traffic lanes or otherwise adjusting the nature of the traffic lanes. The roadway is usually still the same overall width. The width of the lanes within the overall width of the roadway are the target of the changes or adjustments.

Besides referring to these kinds of changes as a road diet, there are some that simply call it “lane reductions” (but that’s not very catchy, is it), and others that refer to it as road re-channelization (a hefty 5-dollar word that makes it sound more scientific).

To those that want to make lane reductions, using the phrase “road diet” is handy since it has a rather positive connotation. We all generally believe that diets are a good thing. Those that oppose road diets are a bit chagrined that the road diet moniker is used and claim that it is a sneaky wording that hides the true intent, consisting of reducing the number of car traffic lanes. In any case, I’ll use herein the phrase road diet.

Depiction Of A Typical Road Diet

Allow me to provide you with an indication of what a road diet might consist of.

Suppose that your town or city has a four-lane road that is known as Main Street.

There are two lanes going in the southbound direction and two other lanes going in the northbound direction. Traffic moves along on this handy thoroughfare. Perhaps this Main Street has been in existence for many years and seemed to serve the needs of the town or city quite well over those many years. It has become a heartened part of the tradition and lore of the place.

But, there are some in the community that have qualms about the layout of Main Street.

It is dangerous for bike riders to ride on Main Street, and even when hugging the curb, there have sadly been periodic instances of car accidents involving wayward automobiles striking kids and adults on bikes. Another concern is that the cars driving on Main Street tend to go faster than the speed limit, often acting like they are driving on a four-lane open highway instead of down a busy street with lots of shops and businesses. There have been many circumstances of pedestrians that almost got hit while trying to cross Main Street.

What to do?

Some would say that this Main Street is primed to go on a road diet.

Here’s what we’ll do.

The total width of Main Street is 44 feet.

There are four lanes of 11 feet each.

Let’s get rid of two of those car traffic lanes, which frees up 22 feet. Since we want to help out the bike riders, let’s use 5 feet respectively on either side of Main Street for a devoted bike lane. In the middle of Main Street, we’ll put a new lane that’s 12 feet wide and allows for making left turns.

Overall, here’s what we originally had for the 44-foot wide Main Street: 

  • 11-foot southbound car lane 
  • 11-foot southbound car lane 
  • 11-foot northbound car lane 
  • 11-foot northbound car lane

Once we put the roadway onto the devised road diet, it would contain this:

  • 5-foot bike-lane southbound 
  • 11-foot car-lane southbound 
  • 12-foot car-lane mix for south/north left-turning traffic 
  • 11-foot car-lane northbound 
  • 5-foot bike-lane northbound

We are still within the original 44 feet of Main Street.

This makes life somewhat easier because if we had wanted to widen Main Street it would have been quite extensive and expensive roadway infrastructure project. We would have had to uproot the sidewalks and the various fire hydrants and light posts. Since we are only changing the lanes within the existing overall width of the road, the effort to make the changes will be a lot less costly and arduous to undertake.

I’m not suggesting that the road diet changes are somehow cost-free or cheap to do.

Having to re-stripe the road and potentially make other modifications to accommodate the new plan can definitively have some substantive costs. If you compare those costs to widening the road or making other more substantive structural changes, on a relative basis the road diet is likely more affordable.

Road Diets Vary

Not all road diets will necessarily stick with the original width of the road.

You can have instances of widening the overall width of the road, even when it is undertaking a so-called diet.

The diet part of things is usually focused on the fact that you are reducing the number of car traffic lanes (or, sometimes reducing the width of the existing car traffic lanes). You then use the freed-up space for other purposes, which might include adding bike lanes, adding a center turn lane, or perhaps widening or adding sidewalks, etc.

For the Main Street example, suppose the changes were indeed made and we have this new and exciting version of Main Street. Bike riders are now less likely (hopefully) to get hit by cars because of the added bike lanes. This might encourage bike riders to use Main Street, more so than they might have otherwise.

Constricting the car traffic to now just one lane in each direction is likely to slow down the cars.

This might deal with the prior aspect that cars were often speeding down Main Street, rising to speeds that were prone to accidents and potentially hitting pedestrians. With the constricted availability of just one lane in each direction, the traffic is perhaps slowed down and will be less apt to drive recklessly.

There are some studies that claim that the proper use of a road diet can reduce car crashes by around 47% and reduce speeding by about 70%.

Those studies also suggest that there might be an increase in bike riders of around 37%. Plus, there might be an increase in pedestrian foot-traffic of around 49%.

It would seem that the use of a road diet is a really good way to keep a roadway in existence and yet redesign and repurpose it to better suit the needs of the community. It can potentially save lives. It can possibly rejuvenate an area – suppose that the bike riders and pedestrians had previously avoided going to the businesses and shops on Main Street due to the car traffic dangers. Now, those bike riders and pedestrians might opt to revisit Main Street and shop there once again.

Road Diets Not All Rosy

Not all is necessarily so rosy on Main Street, though.

If the drop to just one car traffic lane in each direction does constrict traffic flow, it might lead to heavy congestion now on Main Street.

Cars might be snarled all along Main Street, trying to get to their destinations. This heavy traffic might become visually a blight and might also increase noise or odors. It could frustrate car drivers. People might now find themselves taking much longer to drive to wherever they are trying to go and thus burning up more gasoline and wear-and-tear on their cars.

An unintended reaction by the drivers could be that they decide to spillover into nearby neighborhoods.

If Main Street has become a bottleneck, the car drivers might decide to turn onto side streets and weave through whatever streets are adjacent to Main Street. This could impact those streets and endanger those that live on those streets. All of a sudden, a quiet neighborhood that once had only local traffic could now be inundated with cars trying to make their way along on Main Street but desperately seeking an alternative.

This self-diverting of traffic could increase the risks for pedestrians and bike riders in the adjacent neighborhoods.

Whereas maybe Main Street is now safer, it could be that the risks and potential car accidents are merely being shifted into those other streets. You didn’t particularly fix the problem and only pushed it into another area. Sometimes the adjacent neighborhoods now need to react and ask that roadway speed bumps be put in place, along with other traffic restrictions and posted signs to warn car drivers to not carelessly use those adjacent streets as though they are still on Main Street.

Some would argue that another disadvantage of constricting the car traffic involves the potential delaying of first responders to an emergency.

A police car that is trying to quickly get to a crime scene and that uses Main Street might be delayed by the traffic congestion now on Main Street. Likewise, there might be delays to ambulances or fire trucks. This could be another unintended consequence of the road diet.

Another possibility of something amiss could be that cars begin to avoid using Main Street whatsoever.

This certainly reduces the volume of traffic and might aid the use of Main Street for the bike riders and pedestrians. But, it could also lead to less people driving to and visiting the shops and businesses that are lined along Main Street. Those shops and businesses might soon discover that their revenues are drying up due to the road diet.

A somewhat rarer potential issue could be that circumstance of a mass evacuation and the road diet preventing people from readily driving to get out-of-town. If a hurricane is heading in the direction of the town and people are supposed to get out-of-town, perhaps the reduced lanes might slow down all that car traffic and prevent people from fleeing on a timely basis.

As you can hopefully discern, the road diet is a practice that often involves great controversy.

Controversy And Road Diets

Many cities or towns that start toward using a road diet approach will often do so quietly and without much fanfare.

It kind of slips under the radar of the populace. A particular town might decide that they have a roadway they want to put on a road diet. They move forward doing so. Once they are done, all of a sudden, the car traffic that has routinely been using that roadway now becomes quite concerned about what has taken place. People rise up and complain.

It could be that the car traffic that was using that portion of the roadway as a kind of pass-thru and they really didn’t care much about the local aspects per se. If the pass-thru was in a location that does not allow for ready alternatives, it is likely those car drivers are going to be steamed about the changes. The potential political pressure and public backlash can be tremendous.

Plus, even if those impacted are sympathetic to the road diet, they might argue that the particular road diet design was deficient.

As I mentioned earlier, people that sometimes choose to go on a food diet discover that not all food diet programs are necessarily the best for you. It might be that the road diet approach undertaken is not the best choice for the situation at-hand and thus it could be that some that oppose the road diet are primarily opposing the specific implementation of it, and yet still open to a road diet of some alternative design.

One such example of a road diet controversy took place in Santa Monica, California , and I was caught up in the matter as someone that at the time routinely drove through the area in question.

Press coverage encompassed both those in favor of the road diet and those in opposition.

The opposing forces called it a draconian lane reduction, some called it a debacle, some said it was a road diet disaster. There were even efforts to recall some of the politicians that had been involved in the road diet effort.

As you might guess, there was a lot of hand wringing too because some that generally believe in road diets were worried that if this road diet was expunged it might curtail all future road diet initiatives.

Generally, there is much debate on all sides of the road diet approach.

I mentioned earlier that it could be that the road diet might slow down first responders, but this is considered a controversial point and there are some researchers that say it is a false assumption and a myth.

There are numbers and stats to be found on each side of the coin about road diets. It is difficult to compare road diets since they each have their own particular shapes and sizes. A road diet might do well in one jurisdiction and do poorly in another. You cannot usually carte blanche declare a road diet as good or bad, and instead would need to look at the circumstances and situation involved.

Back to my analogy about food diets to the nature of road diets. Are all food (roadway) diets bad? Nope. Are some food (roadway) diets bad? Yes. Is a food (roadway) diet that is good for Joe (city X) necessarily also good for Samantha (city Y)? No. Should we be willing to consider a food (roadway) diet? Yes. Is one food (roadway) diet the same as another? Not usually. And so on.

Road Diets And AI Autonomous Cars

What does this have to do with AI self-driving driverless autonomous cars?

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. One aspect that is considered an “edge” problem involves the nature of road diets as it relates to AI self-driving cars.

It is considered an edge problem since it is not at the core of what most of the automakers and tech firms are focusing on. They pretty much are focused on the rudiments of getting the AI to drive a self-driving car. The use of an AI self-driving car in a road dieted situation is not at the core of the driving task, in their view, and can be dealt with at a later time (it is considered on the corner or edge of the core problem being solved).

For my article about edge problems for AI self-driving cars, see: https://aitrends.com/selfdrivingcars/edge-problems-core-true-self-driving-cars-achieving-last-mile/

I’d like to first clarify and introduce the notion that there are varying levels of AI self-driving cars. The topmost level is considered Level 5. A Level 5 self-driving car is one that is being driven by the AI and there is no human driver involved. For the design of Level 5 self-driving cars, the automakers are even removing the gas pedal, the brake pedal, and steering wheel, since those are contraptions used by human drivers. The Level 5 self-driving car is not being driven by a human and nor is there an expectation that a human driver will be present in the self-driving car. It’s all on the shoulders of the AI to drive the car.

For self-driving cars less than a Level 5, there must be a human driver present in the car. The human driver is currently considered the responsible party for the acts of the car. The AI and the human driver are co-sharing the driving task. In spite of this co-sharing, the human is supposed to remain fully immersed into the driving task and be ready at all times to perform the driving task. I’ve repeatedly warned about the dangers of this co-sharing arrangement and predicted it will produce many untoward results.

For my overall framework about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

For the levels of self-driving cars, see my article: https://aitrends.com/selfdrivingcars/richter-scale-levels-self-driving-cars/

For why AI Level 5 self-driving cars are like a moonshot, see my article: https://aitrends.com/selfdrivingcars/self-driving-car-mother-ai-projects-moonshot/

For the dangers of co-sharing the driving task, see my article: https://aitrends.com/selfdrivingcars/human-back-up-drivers-for-ai-self-driving-cars/

Let’s focus herein on the true Level 5 self-driving car. Much of the comments apply to the less than Level 5 self-driving cars too, but the fully autonomous AI self-driving car will receive the most attention in this discussion.

Here’s the usual steps involved in the AI driving task: 

  • Sensor data collection and interpretation 
  • Sensor fusion 
  • Virtual world model updating 
  • AI action planning 
  • Car controls command issuance

Another key aspect of AI self-driving cars is that they will be driving on our roadways in the midst of human driven cars too. There are some pundits of AI self-driving cars that continually refer to a utopian world in which there are only AI self-driving cars on public roads. Currently there are about 250+ million conventional cars in the United States alone, and those cars are not going to magically disappear or become true Level 5 AI self-driving cars overnight.

Indeed, the use of human driven cars will last for many years, likely many decades, and the advent of AI self-driving cars will occur while there are still human driven cars on the roads. This is a crucial point since this means that the AI of self-driving cars needs to be able to contend with not just other AI self-driving cars, but also contend with human driven cars. It is easy to envision a simplistic and rather unrealistic world in which all AI self-driving cars are politely interacting with each other and being civil about roadway interactions. That’s not what is going to be happening for the foreseeable future. AI self-driving cars and human driven cars will need to be able to cope with each other.

For my article about the grand convergence that has led us to this moment in time, see: https://aitrends.com/selfdrivingcars/grand-convergence-explains-rise-self-driving-cars/

See my article about the ethical dilemmas facing AI self-driving cars: https://aitrends.com/selfdrivingcars/ethically-ambiguous-self-driving-cars/

For potential regulations about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/assessing-federal-regulations-self-driving-cars-house-bill-passed/

For my predictions about AI self-driving cars for the 2020s, 2030s, and 2040s, see my article: https://aitrends.com/selfdrivingcars/gen-z-and-the-fate-of-ai-self-driving-cars/

Road Diets As Aiding Autonomous Cars

Returning to the topic of road diets, let’s consider how road diets are related to AI self-driving cars.

First, some pundits argue that we ought to consider implementing road diets as a means to support the advent of AI self-driving cars.

It is assumed that we will gradually have lots and lots of AI self-driving cars on the roadways.

Furthermore, those AI self-driving cars are likely to be put to use on a non-stop 24×7 basis. You’ll see AI self-driving cars cruising back-and-forth, providing all kinds of ridesharing services for us. You might use an AI self-driving car to take your kids to school and pick them up after classes to drive them home (you won’t need to do the driving, instead just send the AI self-driving car). Or, maybe send your AI self-driving car to pick-up that pizza for dinner. Etc.

For my article about non-stop use, see: https://aitrends.com/selfdrivingcars/non-stop-ai-self-driving-cars-truths-and-consequences/

For my article about ridesharing, see: https://aitrends.com/selfdrivingcars/ridesharing-services-and-ai-self-driving-cars-notably-uber-in-or-uber-out/

For in-car deliveries, see my article: https://aitrends.com/selfdrivingcars/in-car-deliveries-with-ai-self-driving-cars/

Some believe that we should consider blocking off many downtown areas to reduce the amount of car traffic and increase the amount of foot traffic and biking that can occur.

The use of a road diet approach would presumably aid in this notion.

You might then have possibly well-coordinated AI self-driving cars that are streaming back-and-forth throughout these road dieted locations. Those AI self-driving cars are picking up people or goods, and delivering people or goods. They are well-coordinated in that perhaps the use of V2V (vehicle-to-vehicle communications) and V2I (vehicle-to-infrastructure) communications has allowed them to ascertain which ones are going on which roads, and otherwise align their efforts.

This seems sound in theory and nearly Utopian.

As mentioned though, there will be a long-time overlap of human driven cars and AI self-driving cars.

Would both human driven cars and AI self-driving cars be both allowed into these road dieted locations?

If so, the human driven cars would presumably have a more difficult time coordinating their driving activities than would the only-AI driven cars. Also, the variability of how the driving of the human driven cars would occur in the dieted locations is likely higher than the AI self-driving cars (in essence, it would be possible to have the AI self-driving cars driving in a special “road diet mode” while in a road dieted location).

It is likely that these road dieting efforts will encounter many of the same qualms already expressed about today’s road diets. The impact might be that the road congestion generated becomes untenable. It could be that the traffic delays generated become untenable. And so on.

You’ve ordered a pizza delivery to your downtown apartment which is nestled in a road dieted location. Turns out that the AI self-driving car trying to reach you has been delayed in the road diet area due to the volume of cars trying to push through that roadway. You are upset about the delay in getting your pizza!

I realize that the delay in getting a pizza is somewhat silly perhaps. I just used the pizza example to illustrate what might take place. You can substitute the pizza delivery with let’s say medicine being delivered to an elderly person living within the road dieted zone. That obviously ups the stakes in understanding the impact that traffic delays might cause.

Trade-off Aspects Of Implementing Road Diets

Some of the disadvantages of road diets can be likely better controlled via the use of AI self-driving cars.

For example, the spillover effect can be potentially reduced by informing the AI systems of self-driving cars that they are not supposed to try and avoid the road diet by taking side streets. This could be possibly pumped to the AI via the on-board OTA (Over-the-Air) electronic communications, which usually involve providing updates or patches to the self-driving car.

For more about OTA, see my article: https://aitrends.com/selfdrivingcars/air-ota-updating-ai-self-driving-cars/

In general, the aspects of adjusting downtown areas onto a road diet has both its positives and its negatives.

If you could ban human driven cars from those road dieted areas, it might make for a more orderly use of the available car lanes.

Whether humans will put up with being banned from driving in those areas would seem like an open question and one that might generate a lot of public debate and controversy.

Even if you could ban human drivers from those areas, you still need to consider how much car traffic you are anticipating.

I say that because even if you have only AI self-driving cars allowed into a road dieted location, this does not somehow magically overcome the volume and timing of the car traffic axiomatically. You can only get so much water to flow through a pipe of a certain size. The same would be true about the number of car lanes available and the volume of the AI self-driving cars that are being sought to get into and out of the road dieted area at any point in time.

The other major aspect to consider about AI self-driving cars and road diets consists of the specialized driving nature of traversing and using a road dieted location.

There are some AI developers that say there is nothing unusual or new about driving in a road dieted location. In their book, if an AI self-driving car can navigate a normal road, it should be able to do the same when navigating a road dieted location. I consider this to be a head-in-the-sand belief and one that can bode for problems when AI self-driving cars get themselves into such specialized circumstances.

For AI developers and egocentric designs, see my article: https://aitrends.com/selfdrivingcars/egocentric-design-and-ai-self-driving-cars/

We believe that there is more to dealing with a road dieted location than just an everyday driving routine.

Specialized AI Driving Capabilities Needed

Road dieted locations do have a specialized element and therefore merit specialized AI capabilities to properly undertake.

As already mentioned, one aspect would be the driving practices of the AI self-driving car. In a savvy AI system, if there is a roadway bottleneck, the AI will try to find a means to get around the bottleneck. But, in the case of a road diet, it might be that by-design the self-driving cars are being asked to refrain from trying to spillover into nearby neighborhoods. The AI would need to be able to get notified of this driving condition and be able to adjust accordingly.

Another aspect involves whether the road dieted location might have special cut-out areas that are intended for loading and unloading. It is expected that for the safe and efficient delivery of goods and people, there will be various street cut-out areas to allow for loading and unloading, more so than typically is available today. The advent of large volumes of deliveries via AI self-driving cars and ridesharing will increase the need for these specialized zones.

The AI self-driving car needs to be versed in approaching, stopping, and then resuming a car driving journey in these road dieted locations. This must be done with the utmost safety and with the realization that there will likely be a significant presence of both pedestrians and bicyclists.

The AI self-driving car also needs to be ready to cope with the V2V and V2I electronic communications that are likely to be occurring in those road dieted locations, sifting through what might be a voluminous amount of information and coordination aspects.

Another aspect involves the potential for pranking of an AI self-driving car.

It is anticipated that humans might try to “prank” AI self-driving cars, doing so by for example stepping in front of an AI self-driving car to get it to come to a sudden stop. This might be done just for fun or sport, and not due to genuinely needing to get the AI self-driving car to come to a halt. The odds are that this kind of pranking will occur even more so in a road dieted location, due to the higher volume of nearby pedestrians and bike riders. The AI system needs to be versed in dealing with the pranks.

For pranking of AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/pranking-of-ai-self-driving-cars/

If there is human driven car traffic allowed into the road dieted area, the AI needs to be prepared to contend with the variabilities of what those human drivers might do. A savvy AI system needs to be versed in defensive driving techniques overall. Within a road dieted location, there are various specific defensive aspects that the AI should be further have available in its driving capabilities.

For my article about defensive driving for AI self-driving cars, see: https://aitrends.com/selfdrivingcars/art-defensive-driving-key-self-driving-car-success/

Another potential difficulty for an AI self-driving car would be the encountering of a road dieted location for the first time. If the AI self-driving car did not realize beforehand there was a road dieted location on its driving journey, perhaps it is not marked as such on a map or GPS, the AI system needs to detect that a road dieted location exists and that the self-driving car has entered into it. Once having discerned and mapped out the road dieted location, it could potentially add this aspect to its repertoire as part of the Machine Learning capabilities.

For more about machine learning, see my article: https://aitrends.com/ai-insider/machine-learning-benchmarks-and-ai-self-driving-cars/

For ensemble machine learning, see my article: https://aitrends.com/selfdrivingcars/ensemble-machine-learning-for-ai-self-driving-cars/

Road diets, they are coming.

The emergence of AI self-driving cars will likely promote the adoption of lane reductions and road diets.

This should not be done blindly.

A road diet can be a boon to a local area or become a nightmare.

Either way, if a road diet is instituted, the AI of the self-driving car needs to be ready to cope with the particulars of a road dieted location.

It is important to make sure that the AI is beefed-up and not too slim on how to best and safely drive in an area that’s gotten a slenderized road diet.

Copyright 2019 Dr. Lance Eliot

This content is originally posted on AI Trends.