Monday, 4 November 2019

Celebrating Growth at Databricks and 1,000 Employees!

Celebrating 1,000 employees -- Databricks marks the milestone of hiring its 1,000th fullt-time employee.

This November, Databricks hired our 1,000th full-time employee! Founded in Berkeley in 2013, our six co-founders created Databricks to help data teams solve the world’s toughest problems – and since then, we’ve grown tremendously! Not only have we had some major milestones like our Microsoft partnership, resulting in Azure Databricks and the creation of new open source projects like Delta Lake and MLflow, but we have also expanded our employee count and global presence. We now have offices across the world, including London, Amsterdam, Singapore, New York and our headquarters in SF! We are so excited for what’s to come and owe a big thank you to our employees, partners, and customers who have been on this journey with us.

How did we reach 1,000 employees?

At Databricks, we recognize that hiring a new member for our team is a great example of how our employees embody one of our core values “Teamwork makes the dream work!” Providing an excellent candidate experience requires the coordination and collaboration of many teams. With the help of our awesome teammates, Databricks has been able to scale quickly and effectively.

Take a look at all the teams that go into helping a new hire have a great experience!

The path a candidate takes to become a new Brickster -- Databricks employee.

The path a candidate takes to become a new Brickster

How has our company changed from our first year in 2013?

We asked our employees to talk about the biggest changes and growth that they’ve noticed at Databricks from the year that they first started to now, both in the role they play here and the company itself. Learn from their perspective below!

2013: Meet Michael Armbrust, Principal Software Engineer

Michael Armbrust, Principal Software Engineer, Databricks Michael was one of the first engineers hired at Databricks and is a frequent speaker at Spark+AI Summit. He is a committer and PMC member of Apache Spark and the original creator of Spark SQL. He currently leads the team at Databricks that designed and built Structured Streaming and Delta Lake.

Before Databricks, I was a post-doc at Google doing research on building composable optimizers as part of the F1 team. When I joined Databricks, I was really excited to get to put many of the ideas we came up with during that project into production. This lead to the catalyst optimizer that powers Spark SQL today. Even though I studied Databases in grad school, I had never built a “real” one before, and being able to create something like that from scratch at my first job was an amazing opportunity. It was also really gratifying to see hundreds of global contributors from the Apache® Spark community help grow the engine after we released it as open source. Working with such a vibrant community was a big change from working on it on my own for a year.

More recently I got the opportunity to open-source another really cool piece of technology, called “Delta”. Delta started as this proprietary product, that was inspired by a conversation I had with a potential customer at Spark Summit. He was at a large Fortune 100 company, and he wanted to ingest petabytes of data per week into a massive data lake that could be queried in real time by analysts around the world. I knew that a workload like that would overload Spark’s metadata management layer. However, this challenge sparked the idea of creating a scalable transaction log. It was cool to see how many different factors contributed to the start of Delta Lake: the Spark Community, Spark Summit, our sales team, leadership (including our CEO Ali) and the awesome members of the “streamteam” at Databricks. We worked really closely with the customer and in around six months we had gone from an idea to actually running in production! Before long, we decided to share this technology with the world, and earlier this year the Delta Lake open source project was born. While this project is still young, I’m really excited with the momentum so far, and I can’t wait to see where it goes!

2014: Meet Tim Hunter, Technical Lead, Jobs Team

Tim Hunter, Tech Lead, Jobs Team, Databricks Tim started off at Databricks as a Software Engineer on the Machine Learning Team. He did his Ph.D. in the AMPLab, the UC Berkeley lab that created Apache Spark. As part of his research, he wrote Machine Learning algorithms using Spark 0.0.2. Wanting to learn more about the impact of our product, he transferred over to become a Solutions Architect in London, and recently moved to Amsterdam to lead our Jobs Team.

When I started at Databricks, there were 15 employees cramped in tiny office right down the hill from UC Berkeley. Back then, a lot of our work was groundbreaking and done in a semi-stealth session. It was finally unveiled during the first Spark Summit, in what we used to call the “Mother of all demos”. My early work was to make Spark much easier to use with a service called the “chauffeur”, that connects data science to clusters – so people could point and click and not have to worry about the details behind the scenes. After the first two years, I worked on the ML team, with a focus on deep learning, deployment and AI efforts inside Databricks. This is where I looked at how we could add and scale more intricate functions of machine learning like DBUs, image processing, geospatial processing, and graph processing.

One thing that I enjoy about engineering at Databricks is that you’re not just writing code, you’re also thinking globally about how to present your work to users and getting feedback on it. If you also enjoy speaking, like I do, there are a lot of opportunities to give presentations about your work, especially at Meetups or Spark Summit. Since I wanted to understand how the ML system I was building was being used, I asked to do a rotation as a Solutions Architect in our London office. The excitement and the small team reminded me of the original small startup form a few years ago – except that it had the full backing of the U.S. team and a proven, very successful product to sell. There, I helped large European companies put together some ML/AI solutions on top of Databricks in a wide variety of industries such as car manufacturing, drug processing, and chemical processing. After this rotation, I moved to our Amsterdam office to be the Tech Lead of our newly formed Jobs team. It is amazing to see that the service maintained by the jobs team, which sprouted out as a quick hackathon project a few years ago, has developed into an industrial strength system that is pretty much used by every Databricks user! The flexibility that Databricks has given me to explore different offices and roles within the company has really helped me grow as an engineer, and allowed me to see the different areas of growth that we have gone through in our product.

2015: Meet Jen Aman, Senior Event Manager

Jen Aman, Sr. Manager Marketing, Databricks Jen started off as a Marketing Manager and helped Databricks launch our third Spark Summit when it was still only being held in the United States. She is now a Senior Event Manager, and plans and develops the strategy for large scale global events for Databricks, including our annual Spark + AI Summit (Americas and Europe) and company retreat.

I started two weeks before our third Spark Summit, which at the time was only being held in both San Francisco (around 2,000 attendees) and the east coast (around 1,400 attendees). We decided that year that it would be a good time to launch Spark Summit Europe for the first time in Amsterdam. This year, we held our 5th Spark Summit in Europe – which sold out! We eventually decided to only have the Americas Summit in San Francisco and moved from hosting it in hotel ballrooms at the Hilton to Moscone Center (one of the largest convention centers in SF), where we now have around 5,000 attendees. My first year, my core responsibility was to figure out booth duty schedule for everybody. Having started only two weeks before, I had no idea who anyone was and had to look at their pictures to figure out names. Eventually, I took on more responsibility for Summit: owning the agenda process, call for papers (community to submit talks), the marketing and execution of these talks, managing speaker attendance, keynote process, swag, catering, space planning and managing the creative.

The content of our Spark + AI Summit conference has also expanded, outside of just additional keynotes and tracks running at the same time. We now have vertical events – health and life sciences and FinTech, as well as networking events, meetups, tutorials, lightning talks and an advisory bar for questions on Apache Spark™ and Databricks. It’s also been a huge change managing our internal attendance – we only had 54 Databricks employees during the 3rd Spark Summit and in 2019’s Summit, we had 800+ employees! Our external audience has also expanded from mainly Spark enthusiasts to adding on Databricks customers and partners, data professionals networks, and our Women in Unified Analytics program. Seeing the evolution of Spark Summit, in addition to other internal events I help launch, has made it really rewarding to see the impact and growth of our events!

2016: Meet Shelby Ferson, Geo Enterprise Account Executive

Shelby Ferson, Enterprise Account Executive, Databricks Shelby started off at Databricks as a Mid Market Rep, when the sales team was less than 20 people. She has since been promoted to a Commercial Account Executive, and now is an Enterprise Account Executive, helping evangelize Databricks and communicating its value to customers and system integrators, while also helping build out regions around the world. .

I had the exciting opportunity to join Databricks when our team was fairly small, our sales team only had less than 20 people globally at that time! Throughout my time here, it has been rewarding to work with customers that are continuously innovating and seeing how our product has been able to support them and their initiatives over these past three years. It’s also refreshing that with our size, I can still walk down to the engineering floor and have technical conversations and learn more about our product, no matter how busy they are. Everyone is willing to help to get customers on board and work as a team to create the best experience for our customers. It’s a great testament that when sales, product, engineering and customer success, work really closely, amazing things happen!

I’ve also been lucky to be closely supported by our sales leadership, who encouraged me to challenge myself and helped accelerate my growth within Databricks. Our executives invest in building a positive sales team culture and have supported initiatives that I’ve worked (along with the team) to help launch, such as our Women of Databricks events and external events like Databricks’ co-sponsored talk with Tableau. This support also extends to when we build out new regions, and how we emphasize the importance of training and embedding our sales culture to our new teammates. Recently, I had the opportunity to help ramp up and support our sales team in our Australia office, and share all the knowledge I had around our product and company. It’s been a once in a lifetime opportunity to be part of the first sales team here, and now getting to see us expanding our sales team to almost 300 people and expanding in regions all over the world!

2017: Meet Yvette Ramirez, Junior Recruiter

Meet Yvette Ramirez, Junior Recruiter, Databricks Yvette joined Databricks in 2017 as a Recruiting Coordinator in the San Francisco office. After focusing on candidate interviewing and hiring experiences, Yvette supported the Field Engineering, Customer Success, and Professional Services teams as a sourcer for a year. Most recently, she transitioned into a role at the Amsterdam office as a Junior Recruiter to help build the Go-To-Market team in EMEA.

When I started as a Recruiting Coordinator, there were a little less than 250 employees, and the recruiting team was just 9 people. The size, chaos, and growth of the startup world drives you to be scrappy and create organization from the ambiguity. This was the main reason why I moved from Florida to California, for the opportunity to grow with a company like Databricks that I believed was going to be really successful. After just one year, our team helped to more than double the size of the company, reaching 600+ employees. As the growth continued globally into EMEA and Sydney, the size and complexity of our team’s work followed. I saw this opportunity allowing me to expand my skill set, working with the help of many teammates to jump into the Sourcing role, and mentoring the RCs who came afterward.

My most recent move brought me to Amsterdam as a Jr. Recruiter, helping build out our EMEA Go-To-Market teams. With a smaller European recruiting team, it’s really exciting to again get the chance to help build out our offices that are rapidly growing. An influx of people creates the opportunity for more diverse perspectives, which leads to better systems, processes, and ultimately can transform the way we operate and hire. Diversity has always been a passion of mine, having helped our diversity committee at the early stages of its inception. Through events like Lunch & Learns and Women in Analytics events at Spark Summit, we’ve strived to make Databricks a more inclusive environment. Building on that work, we strive to think on a larger scale on how Databricks can become one of the most diverse and inclusive places out there. It has been amazing to grow alongside Databricks, and I’m so excited to see what else we accomplish globally!

2018: Meet Kunal Taneja, Senior Manager, Field Engineering, APJ

Kunal Taneja, Senior Manager, Field Engineering, APJ, Databricks Kunal started off as our first employee in the southern hemisphere as our Sr. Manager of Field Engineering. Within 6 months, he eventually took on building out all of our field engineering teams in APJ. He is now responsible for leading, managing and recruiting a team of field engineers in APJ, who help organizations adopt and use Databricks for driving business value from AI, ML and Unified Analytics.

I joined Databricks in Sydney around employee 370 and was the first hire for our Asia-Pacific Region (APJ). I was paired up with an account representative who started about a month after me, and we were tasked to help grow out our region and the investment in APJ. We then brought in an SVP General Manager for our region based out of Singapore, the hub for our central functions, with the focus of ramping up hiring and building out APJ. Within the first 6 months, we had hired 8 people, and my manager asked if I wanted to help build out our region of Solutions Architects to keep up with the growth of our account representatives. From there, I was asked to lead and grow the APJ Solutions Architect team.

It’s crazy to think about how when we first started, we had only 2 employees and no office for 7 months. Now, approaching 1 year, we had to move offices twice just because of our growth in Australia, and have team members everywhere in the APJ region! Because we’re a smaller office, I’ve really enjoyed that we get to work so closely together, and hang out outside of work at happy hours, lunches and boat cruises! From a region perspective, we’ve also been growing tremendously in Australia, Japan, and India, where the cloud market is booming. We now have 8 Solutions Architects on board, and next year we are expected to at least double. Databricks has offered such a unique opportunity to work with some of the best people in the industry and I’m so excited to see our presence in APJ continue to expand!

2019: Meet Amy Reichandater, Chief People Officer

Amy Reichandater, Chief People Officer, Databricks Amy is our Chief People Officer, and joined Databricks with a strong background in creating highly scalable hiring and retention programs, and driving culture, organization development, and total rewards strategies to support the company’s accelerated global expansion. She has some exciting plans for our growth as a company!

When I joined Databricks in May 2019, I was so excited by the people, market opportunity, growth, and the opportunity to help build an amazing company. As I think about the future of Databricks and what we want to accomplish, my main goal is to create an extraordinary and consistent employee experience. My vision is for all employees, globally, to see Databricks as the most important experience in their career, and as a place where they can bring their best selves to work. Regardless of who they are and where they come from, we want them to feel connected to Databricks, understand our mission and feel empowered to make the company better.

This means that we need to be thoughtful about not only keeping the talent bar high, but also how we can create opportunities for our own teams to grow their career internally as Databricks grows. With doing this, we want to ensure our hiring adds strategic value to the business, and that we evolve the culture and operations in a way that helps us achieve our potential as a company. Our team is really excited about creating extraordinary candidate and employee experiences as we scale, and I’m looking forward to seeing us continue to hire the best talent and build amazing teams around the world!

We are so proud of the team that we have been able to grow with our company and we’re not stopping! Interested? Find your place at Databricks.

--

Try Databricks for free. Get started today.

The post Celebrating Growth at Databricks and 1,000 Employees! appeared first on Databricks.

New Microsoft Azure Data Warehouse Service and Azure Databricks Combine Analytics, BI, and Data Science

In the last two years since it first became available, thousands of companies have adopted Azure Databricks, making it one of the fastest growing data and AI services on Microsoft Azure. Customers now process over 2 exabytes per month with millions of server-hours spinning up every day. All of this is driven by organizations like Electrolux, Shell, and renewables.AI that are using Azure Databricks to process data at massive scale for data science and analytics.

Within this amazing adoption is a specific solution architecture to highlight called the Modern Data Warehouse (MDW). Earlier this year we wrote about the performance and scale benefits of this solution, and part of the pattern’s success has been our close integration to Azure SQL Data Warehouse with a high-performance connector that was jointly engineered to make it fast and easy to move data between the two services.

Three ways Azure Databricks works with Azure Synapse Analytics

Today, Microsoft announced the next evolution of their data warehouse service: Azure Synapse Analytics. This is exciting news and we continue to work closely with Microsoft to integrate with Azure Synapse and bring analytics, business intelligence (BI), and data science together in one solution architecture. Here are three key ways Azure Databricks works with Azure Synapse:

    1. The high-performance connector between Azure Databricks and Azure Synapse will enable fast data transfer between the services, including support for streaming data. This means customers can continue to use Azure Databricks (up to 50x faster than open source Apache Spark) for extract, transform, and load (ETL) workloads to prep and shape data at scale for Azure Synapse.
    2. Azure Data Factory (ADF) supports Azure Databricks in the Mapping Data Flows feature. This offers code-free visual ETL for data preparation and transformation at scale, and now that ADF is part of the Azure Synapse workspace it provides another avenue to access these capabilities.
    3. Azure Synapse and Azure Databricks can run analytics on the same data in Azure Data Lake Storage. This opens even greater opportunities to combine analytics, BI, and data science solutions with a shared data lake across services.

We would love to hear your feedback as you begin using Azure Databricks and Azure Synapse in the next evolution of the Modern Data Warehouse solution architecture.

--

Try Databricks for free. Get started today.

The post New Microsoft Azure Data Warehouse Service and Azure Databricks Combine Analytics, BI, and Data Science appeared first on Databricks.

Sunday, 3 November 2019

Genome sequencing: A solution to India's problem of rare genetic diseases

A genome is a person’s complete set of DNA, including all genes with more than 3 billion DNA base pairs. Genome sequencing will lead to precision medication, instead of clinicians giving drugs based on collective knowledge.

Friday, 1 November 2019

DevOps World | Jenkins World San Francisco in Living Colors

DevOps World | Jenkins World San Francisco was August 12 - 15, 2019. The event was delivered in vivid colors starting with flowing banners hung from street lamp posts to the big screens in breakout rooms, to the expo hall. The energy and enthusiasm in the Moscone convention center made the colors even more vibrant, thanks to the people attending the conference.

Here’s a recap of the conference in pictures:

DevOps World | Jenkins World 2019 - San Francisco

1D5 2235

Keynote - Evolution of the Continuous Delivery Foundation

Tracy Miranda opened the keynote explaining the evolution of the Continuous Delivery Foundation.

Keynote - Evolution of the Continuous Delivery Foundation

1D5 0437

Influencers, Creators, and Members

The influencers, creators, and members of the CD Foundation: Tracy Miranda (far left), Andy Glover (Netflix), Tara Hernandez (Google), Chris Aniszczyk (Linux Foundation), Dave Stanke (Google), Kohsuke Kawaguchi (Jenkins creator), Jayne Groll (DevOps Institute), James Strachan (Jenkins X creator). “We want to help set Jenkins up for success, into the next decade”, Tyler Croy (not in picture).

Influencers, Creators, and Members of the Continuous Delivery Foundation

1D5 1310

Contributor Summit

The inaugural Continuous Delivery Foundation Contributor Summit and it was a full house!

Continous Delivery Foudnation Contributor Summit

IMG 8264

15 Years of Jenkins

A remarkable milestone for the Jenkins project, a celebration of Jenkins turning 15…​cake included!

Fifteen Years of Jenkins

1D5 0614

Bee Diverse Luncheon

Interactive and engaging luncheon celebrating diversity

Bee Diverse Luncheon Entrance

1D5 5576

Bee Diverse Luncheon Leading Voices

1D5 5606

Bee Diverse Luncheon Group Discussions

1D5 5682

Jenkins Contributors and Experts

Jenkins contributors and experts on hand to educate and share lightning talks and provide one on one Jenkins support.

Jenkins Lightning Talks

1D5 3207

Jenkins Experts Answering Questions

1D5 2953

Jenkins Experts Discussing and Helping

IMG 8278

Jenkins Experts Gathered

1D5 3573

DevOps Superheroes

Even though the conference offered endless learning and networking possibilities, and major milestones worth celebrating, I felt the true highlight of the conference was the celebration of each individual, “You”. “You” are the super hero, the driving force behind the incredible innovations to advance technology to where it is today. Here’s celebrating the super heros in all of YOU!

DevOps Superheroes

1D5 2286

Superheroes and the Wookie

1D5 1643

Four Superheroes

1D5 2949

Kohsuke Kawaguchi - Founding Superhero

1D5 3034

A DevOps League of Superheroes

1D5 4067

Crowd of Superheroes

1D5 4243

This party will be coming to Lisbon, Portugal on December 3-5, 2019. We hope to see our EU Jenkins fans at DevOps World | Jenkins World Lisbon. Use JWFOSS for a 30% discount off your pass.

Hope to see you in Lisbon!

7 Ways How AI Will Change Your Workplace

In the next five to ten years, your workplace will look fundamentally different. Thanks to technologies such as artificial intelligence, the internet of things and robotics work as we know it will drastically change. The future of work will come with great opportunities but also with plenty of challenges for organisations. It will require employees and management to adapt and work smarter. AI will augment your jobs, the Internet of Things will provide you with details insights and robotics will replace many jobs.

In the coming decade, your workplace will be datafied and digitalised. Digitalisation refers to the conversion of information into digital format, for example converting music into MP3 files, photos into JPEG, text to HTML and analogue video to YouTube videos. Doing so will increase your available data exponentially. Digitalisation, therefore, means capturing human ideas in digital form for transmission, manipulation, re-use and analysing.

Datafication, on the other hand, refers to turning analogue processes and customer touchpoints into digital processes and digital customer touchpoints. Datafication is the process of making a business data-driven – by transforming social action into quantified data. It involves collecting (new) data from various sources and processes using connected devices or creating detailed customer profiles.

Datafying your ...


Read More on Datafloq

The 4 Up and Coming Data Analytics Trends That Will Define Business for 2020

It is hard to believe that in just a few months we will be heading into an entirely new decade! 2010 to 2019 was certainly a period of incredible development in terms of technology – and there is little doubt that 2020 will continue this trend.

Many experts agree that this next year will be fueled even more by data and analytics. According to Statista, the forecast for Big Data market size is predicted to grow from $49 billion to $56 billion, with steady increases expected every year afterwards.



So, in what specific ways will this market grow?

1. Deep Data Will Become More Accessible than Ever

Analytical platforms are certainly making it easier for every company to obtain and track key data points. As of 2017, more than half of all companies were utilizing data analysis in some capacity.

However, many of these systems are simply running the numbers and churning out reports or charts – of which don’t necessarily make much sense, nor do they always offer practical applications.

In a report from NewVantage, 85.5% of businesses claimed that implementing data throughout their organization was a top priority. But, 63% failed to do so, often due to the fact that there was a lack ...


Read More on Datafloq

Data Science – 8 Powerful Applications

Data science is one of the most exciting emerging fields. As we will see throughout the course of this article it is increasingly becoming an important part of a company’s future. First, we will explain, in simple terms, how data science works and how it can be applied to unstructured data sets. We will also look at how data science is already is widely in use. From medical research to dynamic pricing and driverless cars, data science is increasingly leading the way.



What is Data Science?

Data science is a term given to the practice of analysing raw data to discover any hidden patterns. Various applications and tools such as machine learning and sophisticated algorithms are all used in this process. Unlike other forms of data analysis, it can be applied to both structured and unstructured data. Data science is a more in-depth, detailed way of analysing data than data analytics.

Put simply data analysts use processing history to explain data.

Data scientists employ exploratory analysis and sophisticated tools to uncover new insights and predict future events. This makes it useful for predictive causal analysis. These are models that predict the possibility of a certain event occurring in the future. It is also useful for prescriptive analytics, intelligent models capable of making their own ...


Read More on Datafloq

Sensors and Machine Learning: Glucose Monitoring with An AI Edge

Medtronic’s mission is to alleviate pain, restore health, and extend life through the application of biomedical engineering, explains Elaine Gee, PhD, Senior Principal Algorithm Engineer specializing in Artificial Intelligence at Medtronic. It’s a mission Gee is well equipped for. With over 15 years’ experience in modeling, bioinformatics, and engineering, she drives machine learning algorithm development and analytics to support next-generation medical devices for diabetes management.

On behalf of AI Trends, Ben Lakin, from Cambridge Innovation Institute, sat down with Gee to discuss her most recent focus: algorithm development related to glucose sensing to improve the accuracy and performance of continuous glucose monitoring devices, also known as CGMs.

Elaine Gee, PhD, Senior Principal Algorithm Engineer at Medtronic

Editor’s Note: Gee will be giving a featured presentation on Advancing Continuous Glucose Monitoring Sensor Development with Machine Learning at Sensors Summit in San Diego, December 10-12. This conversation has been edited for length and clarity.

AI Trends: CGMs are a hot topic in the medical device field; they’ve revolutionized monitoring for diabetes patients. Where do Medtronic’s CGM algorithms operate (edge computing, smartphone, the Cloud)? How is Medtronic utilizing new technology to benefit their CGM users?

Medtronic transforms raw user data into personalized insights by various computing methods, from computing on device to computing online. The most advanced sensor from Medtronic is the Guardian Sensor 3. This sensor powers the Guardian Connect CGM. An algorithm performs computations on the device to provide readings every five minutes of user glucose levels and sensor health. By performing analytics in physical proximity to the source where data is generated, reliable real-time insights are available to the user.

Users of the Guardian Connect standalone CGM also receive access to the personal Sugar.IQ diabetes smart assistant. This combines artificial intelligence with diabetes expertise in an app that analyzes daily glucose patterns and factors that affect them, such as food intake or exercise. This informs the user in real-time to encourage around-the-clock glycemic control.

Lastly, users of Medtronic’s MiniMed 670G Insulin Pump System can access the secure web-based CareLink software. This software processes sensor and pump information uploaded by the user to generate personalized insights by highlighting data trends and associating relationships between glucose levels and insulin usage, carbohydrate intake, exercise, and medication to help identify patterns for better glycemic control. The CareLink system also allows collaboration between healthcare providers and patients. With the CareLink Pro software, patients may grant access for providers to view and even download data into their EMR.

These devices must meet numerous design requirements to be considered safe and effective. What are some of the greatest engineering or technical challenges in developing successful algorithms?

Multiple elements come into play when developing a medical device that can pass regulatory standards to get to market. The number one challenge in algorithm design is having enough data. Algorithm development relies on data in every aspect of the development cycle, from the very start with prototyping, to training and optimizing the algorithm, through to testing and validation. Algorithm development begins with proper experimental design to ensure that the data collected to train the algorithm contains the necessary information. Machine learning algorithms rely on high-quality data with proper information density to support learning complex models in a way that avoids over or under fitting. The goal is to create a machine learning model that generalizes well to data not encountered during training to ensure safe use on the market. All these steps, in addition to the regulatory review process, are necessary to verify that the trained algorithm meets the rigorous requirements for safe and effective use in medical device applications.

With regards to future upgrades, what do you think are advantages to pursuing software upgrades or algorithm changes, as opposed to hardware improvements? Is there any difference when it comes to regulatory oversight?

With respect to regulatory oversight, both software and hardware upgrades will trigger a regulatory review. As far as product development, each type of upgrade can improve the CGMs performance and usability. Typically, the R&D process for developing a new algorithm may have a faster turnaround time as compared to the process for supporting a hardware update for various reasons.

First, algorithm development requires large datasets to support training and learning the model. In the cases where training data is unavailable, clinical trial studies are required to collect new data for algorithm development. In this case, this time intensive process can add to the algorithm development timeline as clinical trials must be arranged, subjects enrolled, and data collected.

In terms of algorithm design and experimentation, these steps are done in silico and can be sped up by taking advantage of parallel processing on scalable compute infrastructure. Parallel processing can explore changes to the algorithmic architecture and the parameter space more efficiently to reduce the development time.

Compare this process to making changes to CGM sensor chemistry or other hardware elements. Hardware upgrades require both in vitro and in vivo testing to evaluate performance. In addition, there are manufacturing considerations to address, such as sensor assembly, packaging, sterilization, and shelf life.

Implementing and commercializing new algorithms for CGMs can be more straightforward than hardware changes in some cases, but nonetheless creating a new algorithm requires significant time in development, testing, and validation to ensure the analytic performance supports its medical device use claim and to ensure safe and effective use.

A topic that’s frequently in the news is cybersecurity. With even medical devices at risk from the unscrupulous, how is Medtronic protecting patient data?

I’m glad you asked this, because protecting patient data is a high priority for Medtronic. Cybersecurity is a critical element we incorporate into our development process from the beginning. We take necessary measures to ensure patient data is secure during generation, in transit, and when being stored. Our internal processes guide us to use industry best practices and state-of-the-art approaches. We work with security experts to confirm that our products meet our rigorous criteria, and we closely monitor our products and systems to ensure ongoing protection through the lifecycle of the device.

I’m curious about the possibility of wearable sensors being applied to other conditions. Do you think there are other conditions or medical applications that can benefit from continuous monitoring and real-time data analysis?

Wearable sensors have become accessible and widespread. Real-time data is collected in various forms, whether from direct-to-consumer activity trackers like Fitbits, Apple Watches, and smartphones, or collected from continuous monitoring medical devices like the CGM that we have here at Medtronic Diabetes. With the advancement of wearable sensor technology, consumers and healthcare providers have access to a diverse range of continuous physiological signals. This is exciting because it creates opportunities for new applications of real-time data analysis that could benefit users living with chronic conditions. For example, in the case of diabetes, CGMs and insulin pumps have given users greater control in diabetes management by providing actionable, real-time information. Continued development of sophisticated decision support systems that leverage diverse input signals could give users living with chronic conditions even greater control over their health by generating more actionable insights in real-time.

Are there any new algorithms or hardware technologies on your radar? As an industry scientist, what are you excited about right now?

There are so many technological advances I could list, but given our recent discussion on edge computing I will highlight advances in computational processing by AI microchips. Machine learning algorithms rely on complex computations, placing high demands on the compute, memory, and storage of the hardware chipset. For example, deep learning algorithms frequently perform high-dimensional matrix calculations and thus require large on-board memory to support operations with the hyperparameters and model variables as well as high computing density to support calculations for rapid model inference. Typically, portable wearable devices are battery powered putting constraints on power usage. In order to advance the capabilities of AI-driven wearable devices, algorithms must become more efficient in footprint and power usage, and/or chips must become more powerful and efficient without driving up cost. Algorithms that are more efficient at complex modeling can operate on a smaller hardware footprint, while a faster and cheaper microchip would support high-density calculations at reduced power and cost. Research and development in the area of low-powered AI chip design is exciting because advancements in microchip technology on the hardware side will expand the possibilities for heavier machine learning workloads on the software side.

Learn more about AI at Medtronic.

Collaborative Robots with AI are the Focus of MIT Researcher Julie Shah

By John P. Desmond, AI Trends Editor

“I work on making robots better teammates,” Julie Shah told attendees at 2019 AI World Conference & Expo in Boston. The MIT Associate Professor of Aeronautics and Astronautics described her work in a keynote talk entitled “Enhancing Human Capability with Intelligent Machine Teammates.” She said, “We’re trying to enhance human capability rather than replace humans.”

Also Associate Dean of Social and Ethical Responsibilities of Computing at MIT, Shah directs the Interactive Robotics Group, which designs collaborative robot teammates that aim to enhance human capability. Prior to joining the MIT faculty, she worked at Boeing Research and Technology on robotics applications for aerospace manufacturing.

About 1.6 million robots are operating in the world today, Shah said. However, the number of what could be defined as robots in US homes today number 30 million, including Roombas but not including Alexa, smart homes or autonomous cars. She wondered why our perception is of so many fewer robots in the home than there actually are. “The Roomba does not make an impact on our consciousness,” she said. “That’s a limitation.” She showed a photo of Marty, the robot in many Boston-area Stop and Shop supermarkets, which alerts shoppers when it thinks it sees a spill in the aisle.

Julie Shah, Associate Professor of Aeronautics and Astronautics, MIT

Many industrial robots are deployed in the auto industry, however, 50% of the auto is still assembled by human hand. Some robots have to be physically separated from human workers for safety reasons. Shah concentrates on having humans and robots working in close proximity, and the robot being smart enough to not to cause harm. She prefers the term “mostly not dangerous” robots instead of “safe”, and believes they can augment human capability.

Robots Restricted on San Francisco Sidewalks

Safety is a challenge, Shah said. She recounted how robotic devices deployed on the sidewalks in San Francisco drew complaints from older folks who found it less safe to walk. In 2017, San Francisco was seeing delivery bots and even a security bot to watch sleeping animals at the SPCA shelter. The city Board of Supervisors voted in restrictions that effectively shut down the sidewalk robots, then were relaxed in mid-2018 to allow some continued testing.

“We have a new world where the robots are in our homes and on our streets but are not collaborating with us; they are coexisting with us,” Shah said. “If we don’t get this right, we cut off an innovation path.”

She is studying what makes a good team member. In aeronautics, the pilot works closely with the co-pilot. She proposed a fitting illustration for the Boston audience, “What makes Tom Brady a great team member? He sees and he knows.”

Labor and Delivery Head Floor Nurse Being Assisted by Research Robot

In order to build a seeing and knowing teammate, Shah’s group develops a behavioral model of how people will behave and move in a physical space. Her team studied the work of a nurse running a labor and delivery floor of a hospital, deciding which patients go to what rooms and what staff will attend to them. “The job makes her like an air traffic controller. And she does the job without any decision support. The complexity of the work the cognitive burden, is comparable to that of an air traffic controller,” Shah said.

Her team worked on AI within a robot that would assist the nurse. “We need to be able to learn, like a human apprentice, with very little data.” She called it a “cognitively inspired” approach to learning, and “multi-tiered” decision-making.

“Humans learn by ‘differencing,’ by pair-wide comparisons,” as in asking what was different about what you did yesterday from today. Her team is developing apprentice learning strategies for addressing a wide set of practices.

She showed a video of the robot developed by her team working beside the labor and delivery floor nurse, looking at the white board with patient information. The nurse asked the robot, “What’s a good decision?” and the robot made a recommendation to put a patient in a certain room and have a specific staff person (doctor or nurse) assigned to the room. Then the nurse asked, “What’s a bad decision?” and the robot would respond with a different idea. “I agree with you,” the nurse said, reinforcing what she wanted to do in the first place.

The nurses and physicians are accepting the robot’s suggestion 90% of the time, Shah said. “So we have a high-quality model,” she said, noting that very few people can do the job of the head floor nurse well, and 80% of medical errors with serious consequences are the result of human error.

“So the ability of the system frees up the cognitive capacity of the nurses,” Shah said. “The system can effectively model human capability and how the person will make a decision. Then it plans its own action, to give the person space or to get more involved.”

“It’s not enough for the robot to take an action. We need communication between the person and the machines. That has an enormous impact. Asking questions and then commanding another person are hallmarks of ineffective communication. It’s better to make a suggestion,” said Shah.

In response to an audience question, Shah said “AI applied to robotics opens up fundamentally new opportunities.” The projections for growth in collaborative robots in the next five to seven years she called “remarkable.”  She said, “AI is touching every aspect or every sector in our economy.”

Learn more at MIT’s Interactive Robots site.

AI World Conference & Expo Hosts Its First AI Data Science Hackathon

By Benjamin Ross

BOSTON—The AI World Conference & Expo featured its first ever AI Data Science Hackathon last week, which gave data scientists and developers from across the ecosystem the opportunity to solve real-world data challenges in applying artificial intelligence (AI) and machine learning.

Over the span of three days, teams worked to improve pipelines, datasets, tools, and other projects from a wide range of disciplines.

Two teams gave reports on their work to the AI World audience, one team focused on strategic planning powered by AI in the cloud, and the other working on a fractal AI model for versatility, speed, and efficiency.

Team one—designated “AI-Driven Strategy”—discussed the benefits of strategic planning for businesses with the assistance of AI. “I believe two things about strategic planning in most organizations,” the team’s leader said during their report out. “[First,] it has the potential to give an organization a powerful competitive advantage, if it’s performed competently. Second, it is usually not performed competently.”

AI-Driven Strategy’s solution is to apply AI and machine learning to automate the entire strategic planning function for every organization in the world. Of course, such an ambitious goal would take years to develop, the team leader says, and would reach a scale comparable to the Manhattan Project.

Ideally, the team’s strategy would include applying AI during the initial stage of strategic planning, which includes collecting data from four key domains: the resources and competencies within the organization; the targeted markets and customers; the industries and competitors potentially preventing the organization from reaching said markets and customers; and regulations, economics, demographics, and technologies that keep the organization on track.

For the Hackathon, the team attempted to develop algorithms that would take data dealing with markets and customers as they deal with healthcare organizations and create actionable insights.

“What we’re talking about here is a tool that’s going to bring strategic planning from the 19th century to the 21st century,” the team leader said. “This is a removal of the strategic planning model that would occur maybe once a year, where the top executives of a firm… would receive some data prepared by staff members and use their great wisdom acquired through ‘years of experience’ to make some decisions about strategic plans… Some attempt may have been made to implement them, but generally it wouldn’t work well, and as a result the organization would become disenchanted with the whole process.

“What we have in mind will be utterly dynamic, it’ll be happening all the time… Companies will be seeing these data constantly, and they’ll have the opportunity—with guidance from these algorithms—to make changes, to make adjustments, and to gain the competitive advantage that we believe is available through the strategic planning process.”

Hot Topic

Team two relied on an existing neural network architecture to tackle large spatial and time datasets. The architecture—called the Fractal Artificial Intelligence Model (FAIM)—was used by the team to predict the occurrence of forest fires in the U.S., with the end goal being to leverage that data to enable firefighters to take preventative action.

The team was led by FAIM’s co-founders, Jan Gerards and Jeroen Joukes, who told the audience an advantage of FAIM is its ability to analyze any kind of dataset quickly efficiently, and economically with a little amount of hardware needed. In fact, Gerards said, their work during the Hackathon was done entirely on a Raspberry Pi, a credit card-sized, low cost computer.

“We wanted to show the power of [FAIM], and the value propositions it can bring to the table,” said Gerards. “We hope that this can be disruptive in a positive way.”

The team looked at data from the Office of Satellite and Product Operations (OSPO), including longitude, latitude, temperature data, the size of a particular fire, and fire flags.

Joukes pointed out that FAIM works with previously uncollected data, calculating predictions in real time.

“Once the model is initiated, it trains on a historical dataset from scratch—which takes about five to ten seconds—and then it generates a bunch of predictions, stores them in a local database, and repeats,” Joukes said. “This cycle repeats over and over, collecting evidence for the future that you can use for all types of use cases where it is of the essence to make predictions quickly based on newly acquired information.”

While time constraints proved to be a major factor in the small portions of data collected by the team, Gerards reported that there are still lessons that can be learned from their work.

“I think the problems we faced makes this Hackathon—although we didn’t achieve our desired results— poignant because it’s a reminder that our machine learning capabilities and the abilities to enable AI to provide solutions is really constricted by our data,” said Gerards. “[Data] needs to be cleaned, it needs to be properly packaged together. Otherwise the tool is useless.”

Chaff Bugs and AI Autonomous Cars

By Lance Eliot, the AI Trends Insider

In the movie remake of the Thomas Crown Affair, the main character opts to go into an art museum to ostensibly steal a famous work of art, and does so attired in the manner of The Son of Man artwork (a man wearing a bowler hat and an overcoat).

Spoiler alert, he arranges for dozens of other men to come into the museum dressed similarly as he, thus confounding the efforts by the waiting police that had been tipped that he would come there to commit his thievery. By having many men serving as decoys, he pulls off the effort and the police are exasperated at having to check the numerous decoys and yet are unable to nab him (he sneakily changes his clothes).

This ploy was a clever use of deception.

During World War II, there was the invention of chaff, which was also a form of deception.

Radar had just emerged as a means to detect flying airplanes and therefore be able to try and more accurately shoot them down. The radar device would send a signal that would bounce off the airplane and provide a return to the radar device, thus allowing detection of where the airplane was.

It was hypothesized that there might be a means to confuse the radar device by putting something into the air that would seem like an airplane but was not an airplane. At first, the idea was to have something suspended from an airplane or maybe have balloons or parachutes that could contain a material that would bounce back the radar signals.

A flying airplane could potentially release the parachutes or balloons that had the radar reflecting material.

After exploring this notion, it was further advanced by the discovery that pieces of metal foil could be dropped from the airplane and that it was an easier way to create this deception.

The first versions were envisioned as acting in a double-duty fashion, doing so by being the size of a sheet of paper and contain propaganda written on them. This would be providing a twofer (two-for-one), it would confuse the radar and then once landed on the ground it would serve as a propaganda leaflet.

Turns out that the use of strips of aluminum foil were much more effective.

This nixed the propaganda element of the idea. With the strips, you could dump out hundreds or even thousands of the strips all at once, bundled together but intended to float apart from each other once made airborne. These strips will flutter around and the radar would ping off of them. With a cloud of them floating in the air, the radar would be overwhelmed and unable to identify where the airplane was.

Interestingly, this was considered such a significant defensive weapon that neither the Allies nor the Germans were willing to use them after they had each independently discovered the invention.

It was thought that once used, even if used just one time, the other side would then also discover it and be able to use the same.

We have two examples then of the use of decoys for deception purposes.

One is the bowler hat attired character in a movie, the other is the real-world use during World War II.

Of course, there are many other ways in life that you might come across these kinds of ploys.

One such relatively newer such use applies to computer software.

Use Of Deception And Chaff Bugs In Software

Researchers at NYU had posted an innovative research paper about their use of chaff bugs in software (work done by Zhenghao Hu, Yu Hu, and Brendan Dolan-Gavitt).

This reuse of the old “decoys trickery” is an intriguing modern approach to trying to bolster computer security.

Some of you will find it curious.

Some of you will think it genius.

Some of you with think it silly and unworkable.

Here’s the deal.

We already know that computer hackers (in this case, the word “hackers” is being used to imply thieves or hooligans) will often try to find some exploitable aspect of a software program so that they can then get the program to do something to their bidding or otherwise act in an untoward manner.

Famous exploits include being able to force a buffer overflow, which then might allow the program to suddenly get access to areas of memory it normally should not. Or, it might be that the exploit causes the program to come to a halt or crash. This might be desirable by the hacker either to allow them to take some other action because perhaps the program was acting as a guardian, or that it might cause havoc or confusion that the hacker is hoping to stoke.

Indeed, in the news, there was a reveal that certain HP All-in-One printers could be sent a rigged image to the fax machine portion of the printer, and it would cause a buffer overflow that could allow for someone remotely to get the printer to engage in its built-in remote code execution mode.

Once in the remote code execution mode, the nefarious person could get the printer to do other potentially bad things by sending it additional code and commands. Plus, if the printer was behind a firewall and had internal network access, you could possibly sneak into other of the network connected devices too.

Skilled such hackers are continually on the look for exploits in software.

These potential exploits are usually small clips of code that can be turned into aiding the evil doings of the hacker.

Generally, this is highly skilled work to find the exploits and then devise a means to leverage the exploit. I say that because most of the time when you hear in the news about some “hacker” that got into a person’s computer or email, it is something very low-tech such as they guessed that the person used a password like “12345” and the “hacker” simply used that to break-in.

By the way, these simpleton kinds of break-ins of guessing passwords are somewhat demeaning to the highly skilled hackers. If you are a hunter that is trained and has years of experience in how to use a gun and hunt for wild boar, you are pretty irked when someone at a campground puts out some strips of bacon and the wild boar lands in their lap. Irked because those that don’t know about hunting will equate the trained and experienced hunter with the idiot that happened to have the bacon.

In any case, when developing software, well-versed software engineers and programmers should be trying to avoid writing something that can become an exploit.

If you have in-depth knowledge of the programming language you are using, you should already have familiarity with the known potential exploits that you can fall into. Unfortunately, not all programmers are versed in this, or they are so pressured to write the code quickly that they don’t think about the potential for exploits, or they are not aware of how the code will be compiled and run such that it creates an exploit that they would not otherwise have anticipated. And so on.

Let’s get back to the researchers and what they came up with.

They were trying to develop software that would find exploits in software. Indeed, there are various tools that you can use to find potential exploits. The hackers use these tools. Such tools can also be used by those that want to scan their own software and try to find exploits, hopefully doing so before they actually release their software. It would be handy to catch the exploits beforehand, rather than having them arise at a bad time, or allow someone nefarious to find them and use them.

To be able to properly test a tool that seeks to find exploits, you need to have a test-bed of software that has potential exploits and thus you can run your detective tool on that testbed.

Presumably, the tool should be able to find the exploits, which you know are in the test-bed. This helps to then verify that the tool apparently works as hoped for. If the tool cannot find exploits that you know are embedded into the testbed, you’d need to take a closer look at the tool and try to figure out why it missed finding the exploit. This cycle is repeated over and over, until you believe that the tool is catching all of the exploits that you purposely seeded into the testbed.

So, you need to create a test-bed that has a lot of potential exploits. It can be laborious to think of and write a ton of such exploits. You can find many of them online and copy them, but it is still quite a labor intensive process. Therefore, it would be handy to have a tool that would generate exploits or potential exploits which you could then insert into or “inject” into a software test-bed.

With me so far?

Here’s the final twist.

If you had a tool that could create potential exploits, doing so for purposes of creating test exploits to be put into a test-bed of software, you could also consider potentially using the ability to generate potential exploits to create decoys for use in real software.

Think of each of the potential exploits as akin to a strip of foil for the World War II chaff.

Explaining How Chaff Bugs Work

For the WWII chaff, you’d have lots and lots of the strips, so as to overwhelm the enemy radar.

Why not do the same for software by generating say hundreds or maybe even thousands of potential exploits, clips of code, and then embed those clips of code into the software that you are otherwise developing.

This could then serve to trick any hacker that is aiming to look into your code.

They would find tons of these potential exploits.

Now, you’d of course want to make sure that these seeded exploits are non-exploitable.

In other words, you’d be shooting your own foot if you were generating true exploits. You want instead ones that look like the real thing, but that are in fact not exploitable.

The hacker then would be faced with having to find a needle in a haystack, meaning that even if you have a real exploit in your code, presumably done unintentionally and you didn’t catch it beforehand, the hacker now is faced with thousands of potential exploits and the odds of finding the true one is lessened. This would raise the barrier to entry, so to speak, in that the hacker now has to spend an inordinate amount of time and effort to possibly find the true exploit, even if it exists, which it might not.

I realize that the initial reaction is that it seems somewhat surprising, maybe ludicrous, for you to purposely put potential exploits into your code.

Even if they are truly and utterly non-exploitable, it still just seems like you are playing with fire. When you play with fire, you can get burned. That’s the concern for some, namely that if you inject your (hopefully) clean code with hundreds or thousands of potential non-exploitable exploits, it seems like something bad is bound to happen.

This might be similar to the famous line in the movie Ghostbusters when they were cautioned to not cross the streams of their ghostbuster energy guns, since it could cause an obliteration of the entire universe.

You could say that vaccines are dangerous and yet we use them on humans everyday.

A vaccine is a weakened version of the real underlying virus, and you use it to get the human body to react and build-up a defense. The defense then comes to play when a true wild attack of the virus occurs. This analogy though isn’t quite apt for this matter since the non-exploitable exploits aren’t bolstering the code of the software, instead they are there simply as decoys.

These so-called chaff bugs might be an effective kind of decoy.

Would it scare off a hacker looking for exploits?

Maybe yes, maybe no.

On the one hand, if the hacker looked at the overall code and found right away a potential exploit, they might get pretty excited and think it is their lucky day. They might then expend a lot of attention to the found exploit, which presumably if truly un-exploitable then is a waste of their time.

Would they then give up, or would they look for another one?

If they give up, great, the decoy did its thing. If they look for more, they’ll certainly find more because we know that we’ve purposely put a bunch of them in there.

After finding numerous such potential exploits, and after discovering that they are non-exploitable, would then the hacker give up?

This seems likely. It all depends on how important it is to find a potential exploit. It also depends on how “good” the non-exploitable exploits are in terms of looking like a true exploit.

If the hacker can somehow readily figure out which are the injected exploits, they could then readily opt to ignore them. In that sense, we’re back to the software essentially not having any of the decoys in it, in the sense that if the decoys are all readily discoverable, it’s about the same as if you had none at all in your code.

Therefore, the decoys need to be “good” decoys in that they appear to be exploitable exploits and do not readily appear to be non-exploitable exploits.

This can be tricky to achieve. Having an exploit that is non-exploitable can be achieved by doing some relatively simple things to mute or block the exploit’s exploitability, but then it becomes very easy to look at the exploit and know that it is most likely a planted non-exploitable exploit.

In terms of the planting of the non-exploitable exploits, that’s another factor to be considered.

If I put the decoys in very obvious places of the code, the hacker might realize the underlying pattern of where I am putting the decoys. This then allows the hacker to either ignore those seeming exploits or at least realize that after some limited inspection that if it came from that area of the code it is more likely to be a planted one. You’ve got to then find a means to plant or inject the exploits so that their positioning in the code is not a giveaway.

You could consider randomly scattering the non-exploitable exploits throughout the software.

This might not be so good.

On the one hand, the randomness hopefully prevents a hacker from identifying a pattern to where the decoys were planted.

At the same time, it could be that you’ve put an exploit in a part of the code that would not be advisable for it.

Inner Trickery Is Key

Allow me to explain the logic involved.

If the non-exploitable exploits are to be convincing as potentially exploitable exploits, they presumably need to actually do something and cannot be just stubs of code that are blocked off from execution.

A blocked off exploit would upon inspection be an obvious decoy and thus not require much further effort to explore. Remember that we want the presumed hacker to consume a lot of time by examining the decoys.

But, if the non-exploitable exploits are indeed able to execute then we need to be sure that not only don’t they do something that a genuine exploit would do, we also need to be concerned about their execution time and the consumption of computing cycles.

The decoy might be using up expensive computing cycles, doing so needlessly, other than to try and suggest that it is not a decoy. When I say expensive, I am referring to the notion that the computing cycles might be needed for other computational tasks, and so the decoy is robbing those tasks by chewing up cycles (in addition, you could say “expensive” depending upon what the cost is for your computing cycles).

The decoy might use up both computer cycles and also memory.

Memory could also be a limited resource and therefore the decoy is using it up, solely for the purposes of trying to throw off a potential interloper. As such, however it is that the non-exploitable exploit works, it must appear to be a true exploit, and yet not do any harm, and also minimize consumption of precious computational cycles and computer memory. This is a bit of a tall order as to having non-exploitable exploits that can be such alluring decoys at the minimal overall cost possible.

We could opt to have let’s say simple decoys and super decoys.

The simple decoys are not very convincing and are readily detectable, while the super decoys are complex and difficult to detect as a decoy.

Then, upon seeding of the source code, we might use a mixture of both the simple decoys and the super decoys.

As mentioned before, though, it is important to refrain from putting any of the more time consuming or memory consuming decoys into areas of the code that might be especially adversely impacted. If there’s a routine in the code that needs to run fast and tightly, putting a decoy into the middle of it would likely be unwise and detrimental.

AI Autonomous Cars And Chaff Bugs

What does this have to do with AI self-driving driverless autonomous cars?

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. As part of that effort, we are also exploring ways to try and protect the AI software, particularly once it is on-board a self-driving car and could potentially be hacked by someone with untoward intentions.

I’ve previously discussed that there are many trying to steal the secrets of AI systems for self-driving cars, see my article: https://aitrends.com/selfdrivingcars/stealing-secrets-about-ai-self-driving-cars/

I’ve also discussed that there are numerous computer security concerns underlying the running of the AI of a self-driving car, see my article: https://aitrends.com/ai-insider/ai-deep-learning-backdoor-security-holes-self-driving-cars-detection-prevention/

One approach to help make the AI software harder to figure out for an interloper involves making use of code obfuscation. This is a method in which you purposely make the source code difficult to logically comprehend.

See my article about code obfuscation: https://aitrends.com/selfdrivingcars/code-obfuscation-for-ai-self-driving-cars/

Another possibility of a means to try and undermine an interloper would be to consider using chaff bugs in the AI software.

This has advantages and disadvantages.

It has the potential to boost the security by a security-by-deception approach and might discourage hackers that are trying to delve into the system. A significant disadvantage would be whether the decoys would possibly undermine the system due to the real-time nature of the system. The AI needs to work under tight time constraints and must be making computational aspects that ultimately are controlling a moving car and for which the “decisions” made by the software are of a life-or-death nature.

A decoy that placed in the wrong spot of the code and for which chews up on-board cycles could put the AI and humans at risk.

Consider that these are the major tasks of the AI for a self-driving car:

  • Sensor data collection and interpretation
  • Sensor fusion
  • Virtual world model updating
  • AI action plan updating
  • Car controls command issuance

See my article about my framework for AI self-driving cars for more details about these tasks: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

Where would it be “safe” to put the decoys?

Safe in the sense that the execution of the decoy does not delay or interfere with the otherwise normal real-time operation of the system.

It’s dicey wherever you might be thinking to place the decoys.

We also need to consider the determination of the hacker.

If a hacker happens upon some software of any kind, they might be curious to try and hack it, and give up if they aren’t sure whether the software itself does something of enough significance that it is worth their effort to continue trying to crack it. In the case of the AI software for a self-driving car, there is a lot of incentive to want to crack into the code, and so even if the decoys are present, and even if it becomes apparent to the hacker, and even if the hacker is somewhat discouraged, it would seem less likely they’d give up since the prize in the end has such high value.

Anyway, one of the members of our AI development team is taking a closer look at this potential use of chaff bugs. My view is that the team members are able to set aside a small percentage of their time toward innovation projects that might or might not lead to something of utility. It is becoming gradually somewhat popular as a technique among high-tech firms to allow their developers to have some “fun” time on projects of their own choosing. This boosts morale, gives them a break from their other duties, and might just land upon a goldmine.

Machine Learning And Chaff Bugs

One approach that we’re exploring is whether Machine Learning (ML) can be used to aid in figuring out how to generate the non-exploitable exploits and also to make those decoys as realistically appearing to be integral to the code as we can get.

By analyzing the style of the existing source code, the ML tries to take templates of non-exploitable exploits and see if they can be “personalized” as befits the source code.

This would make those decoys even more convincing.

For more about Machine Learning and AI self-driving cars, see my article: https://aitrends.com/ai-insider/machine-learning-benchmarks-and-ai-self-driving-cars/

At an industry conference I mentioned the chaff bugs work, and I was asked about whether to hide them or whether to make them more obvious in some respects.

The idea is that if you hide them well, the hacker might not realize they are faced with a situation of having to pore through purposely seeded non-exploitable exploits and so blindly just plow away and use up a lot of their effort needlessly.

On the other hand, if you make them more apparent, at least some of them, it might be a kind of warning to the hacker that they are faced with software that has gone to the trouble to make it very hard to find true exploits. You might consider this equivalent to putting a sign outside of your house that says the house is protected by burglar alarms. The sign alone might scare off a lot of potential intruders, even whether you have put in place the decoys or not (some people get a burglar alarm sign and put it on their house as merely a scare tactic).

For AI software that runs a self-driving car, I’d vote that we all ought to be making it as hard to crack into as we can.

Conclusion

The auto makers and tech firms aren’t putting as much attention to the security aspects as they perhaps should, since right now the AI self-driving cars are pretty much kept in their hands as they do testing and trial runs. Once true AI self-driving cars are being sold openly, the chances for the hackers to spend whatever amount of time they want to crack into the system goes up.

We need to prepare for that eventually. If AI self-driving cars become prevalent, and yet they get hacked, it’s going to be bad times for everyone, the auto makers, the tech firms, and the public at large.

Chaff bugs and whatever other novel ideas arise, we’re going to be taking a look and be kicking the tires to see if they’ll be viable as a means to protect the AI systems of self-driving cars.

Copyright 2019 Dr. Lance Eliot

This content is originally posted on AI Trends.

What’s Next For Robotics: In The Field, Inferencing On The Edge

By Allison Proffitt

Robots are a key application for AI and in addition to an excellent plenary talk by Julie Shah of MIT, a whole track was dedicated to AI in robotics applications. Dan Kara, VP of robotics and intelligent systems for WTWH Media, outlined some of the challenges in building robots—not chatbots, he clarified, but robots that act in the physical world. “It seems like every year it’s just around the corner,” he said, but this year the tailwinds are picking up.

Robotics is the foundation for much of our work thus far in artificial intelligence and machine learning, Kara argued. “It’s only been fairly recently that you’ve started getting artificial intelligence or machine learning moving off into different labs,” he said. “At one time, they were considered the same thing because that’s where the work was going.” Early research in facial recognition, accelerometers, natural language processing, very small cameras, and more “came out of work done in robotics labs. They’re naturally synergistic,” he said.

In the past decade, AI and machine learning have exploded, those advances in cognitive capabilities, IoT, data at scale, and ubiquitous connectivity are funneling back into robotics, Kara said, creating pathways for robots that “think, sense, and act.”

We are seeing a shift in robotics over the past five years from an emphasis on hardware to an emphasis on software. “If you ask, for example, the people who run iRobot, most of their engineers—over three-quarters of them—are software engineers. You’re seeing more and more and more of that occur throughout the world in the leading robotics centers.” In journal articles, we see the same trend: an explosion in work dedicated to the intersection of machine learning and robotics, he observed.

Dan Kara, VP of Robotics and Intelligent Systems, WTWH Media

Robots in the physical world that are doing tasks such as grasping, manipulation, autonomous navigation and localization need reinforcement learning, Kara said, as opposed to supervised, unsupervised, or semi-supervised learning that may be more appropriate for other applications. “There’s a greater emphasis on learning that is happening in real time and the software that’s needed to support that,” Kara explained.

In fact, Kara explained, the de facto standard is becoming Robot Operating Systems (ROS) linking to cloud services. Amazon launched AWS RoboMaker last November: a cloud extension for the ROS with a development environment, simulation tools, and fleet management. RoboMaker offer extensions from the ROS into Amazon’s backend service packages, for example, Amazon Polly for language generation. These sorts of packages could easily include facial recognition and object recognition, but they haven’t yet ventured into manipulation and grasping or navigation. “I suspect that will be coming for this particular product,” Kara predicted.

Amazon isn’t alone. Through Microsoft’s Visual Studio you can access ROS nodes to get capabilities, while feeding off Azure for natural language processing, object recognition, fleet management and other tools. Facebook has a product cloud called SciRobotics, which they developed with Carnegie Mellon University, that is home to these packages for ROS use. Google has also teased a product: the Google Cloud Robotics Platform.

A Model of Programmed Teamwork

“We’ve moved from an era where we’re focusing on single robotic systems for use to now multi-robot systems and how they work together in operation,” Kara said.

While it may look like a fleet of warehouse robots are cooperating in an intricate dance, we actually aren’t there yet. It’s all just obstacle avoidance, explained Michael Franklin, assistant professor, College of Computing, Kennesaw State University. There is no communication at all about what the obstacles are, he pointed out. While the robots are programmed not to run into each other, they have no idea what they are avoiding, and they are not working in concert.

Multi-agent, multi-team situations are complicated, Franklin argued—with a quick history lesson on the Battle of Waterloo as evidence.

AI—in robots or elsewhere—does not strategize, Franklin said. It is reactive and always maximizes its own mission. We haven’t yet built AI that understands teamwork, he argued. In teams, individuals sacrifice themselves for the greater goal.

He proposed a hierarchical agent-based model, with intelligent agents at the edge of the field. Each agent has access to knowledge gleaned from teammates; above that there are policies in place for the task; above that an overarching strategy, and finally, intelligence. But reasoning doesn’t only move top-down. Each agent is an intelligent actor that feeds data back up the model. Even if communications are cut off, the edge agents can carry on with the last, best data.

It’s a vision we are still some distance from realizing, but Kara agrees that intelligence on the edge is the future of robotics research and development. Robots—as highly-sensored physical devices—are acting as edge hubs collecting feedback from other sensors, consolidating that data, and sending it on.

“If you talk to some of the leaders in the cloud/AI or cloud/machine learning infrastructure players—the Googles of the world or the Microsofts of the world—they consider robotic systems to be just hyper-sensored, hyper-intelligent edge devices,” Kara said.

There’s an emphasis now on edge inferencing, and Google is researching federated learning. “It all ties into this notion that we want to actually have the inferencing not done in the cloud, but in fact done on the device itself. You see specialized processors coming from Google and from NVIDIA and QUALCOMM and a variety of other players to emphasize that,” he said.

This will be all the more useful as robots emerge from cages in warehouses. “What about the other 99% of the world outside of buildings, where you’re dealing with sparse data, much more experience driven, combining different types of modalities?” Kara asked.

While most of the emphasis thus far has been on business intelligence, post facto, field robotics will require different types of learning and intelligence. “The types of systems that will add the most value are ones that can reduce the time and response between when something happens and the time to react.” Field robotics is where the big investments are, he observed.

New models of learning—federated learning, emergent learning and continual learning—will be needed to train systems that exist in the field and new classes of hardware and software will be needed to support inferencing on the edge.

Learn more at WTWH Media.

Thursday, 31 October 2019

Solving the Challenge of Big Data Cloud Migration with WANdisco, Databricks and Delta Lake

Migrating from Hadoop on-premises to the cloud has been a common theme in recent Databricks blog posts and conference sessions. They’ve identified key considerations, highlighted partnerships and described solutions for moving and streaming data to the cloud with governance and other controls, and compared the runtime environments offered between Hadoop and Databricks to highlight the benefits of the Databricks Unified Data Analytics Platform.

Challenges for Hadoop users when moving to the cloud

WANdisco has partnered with Databricks to solve many of the challenges for large-scale Hadoop migrations. A particular challenge for organizations that have adopted Hadoop at scale is the traditional problem of data gravity. Because their applications assume ready, local, fast access to an on-premises data lake built on HDFS, building applications away from that data becomes difficult, because it requires building additional workflows to manually copy or access data from the on-premises Hadoop data lake.

This problem is exacerbated by an order of magnitude if those on-premises data sets continue to change, because the workflows to move data between environments add a layer of complexity, and don’t handle changing data easily.

While the cloud brings efficiencies for data lakes there remains concerns about the reliability and the consistency of the data. Data Lakes typically have multiple data pipelines reading and writing data concurrently, and data engineers have to go through a tedious process to ensure data integrity, due to the lack of transactions.

Hadoop migration with Databricks and WANdisco

The Databricks and WANdisco partnership solves these challenges, by providing full read and write access to changing data lakes at scale during migration between on-premises systems and Databricks in the cloud. This solution is called LiveAnalytics, and it takes advantage of WANdisco’s platform to migrate and replicate the largest Hadoop datasets to Databricks and Delta Lake. WANdisco makes it possible to migrate data at scale, even while those data sets continue to be modified, using a novel distributed coordination engine to maintain data consistency between Hive data and Delta Lake tables.

LiveAnalytics migrates and replicates the largest Hadoop datasets to Databricks and Delta Lake

WANdisco’s architecture and consensus-based approach is the key to this capability. It allows migration without disruption, application downtime or data loss, and opens up the benefits of applying Databricks to the largest of data lakes that were previously difficult to bring to the cloud.

Because WANdisco LiveAnalytics provides direct support of Delta Lake and Databricks along with common Hadoop platforms, it provides a compelling solution to bringing your on-premises Hadoop data to Databricks without impacting your ability to continue to use Hadoop while migration is in process.

WANdisco’s architecture allows migration from Hadoop to the cloud without disruption, application downtime or data loss

You can take advantage of WANdisco’s technology today to help bring your Hadoop data lake to Databricks, with native support of common Hadoop platforms on-premises and for Databricks and Delta Lake on Azure or AWS.

Related Resources

--

Try Databricks for free. Get started today.

The post Solving the Challenge of Big Data Cloud Migration with WANdisco, Databricks and Delta Lake appeared first on Databricks.