Thursday, 30 January 2020

Google Summer of Code 2019 Report

Google Summer of Code is much more than a summer internship program, it is a year-round effort for the organization and some community members. Now, after the DevOps World | Jenkins World conference in Lisbon and final retrospective meetings, we can say that GSoC 2019 is officially over. We would like to start this by thanking all participants: students, mentors, subject matter experts and all other contributors who proposed project ideas, participated in student selection, in community bonding and in further discussions and reviews. Google Summer of Code is a major effort which would not be possible without the active participation of the Jenkins community.

In this blogpost we would like to share the results and our experience from the previous year.

Results

Highlights

Project details

We held the final presentations as Jenkins Online Meetups in late August and Google published the results on Sept 3rd. The final presentations can be found here: Part 1, Part 2, Part 3. We also presented the 2019 Jenkins GSoC report at the DevOps World | Jenkins World San Francisco and at the DevOps World | Jenkins World 2019 Lisbon conferences.

In the following sections, we present a brief summary of each project, links to the coding phase 3 presentations, and to the final products.

Role Strategy Plugin Performance Improvements

Role Strategy Plugin is one of the most widely used authorization plugins for Jenkins, but it has never been famous for performance due to architecture issues and regular expression checks for project roles. Abhyudaya Sharma was working on this project together with hist mentors: Oleg Nenashev, Runze Xia and Supun Wanniarachchi. He started the project from creating a new Micro-benchmarking Framework for Jenkins Plugins based on JMH, created benchmarks and achieved a 3501% improvement on some real-world scenarios. Then he went further and created a new Folder-based Authorization Strategy Plugin which offers even better performance for Jenkins instances where permissions are scoped to folders. During his project Abhyudaya also fixed the Jenkins Configuration-as-Code support in Role Strategy and contributed several improvements and fixes to the JCasC Plugin itself.

Role strategy performance improvements

Plugins Installation Manager CLI Tool/Library

Natasha Stopa was working on a new CLI tool for plugin management, which should unify features available in other tools like install-plugins.sh in Docker images. It also introduced many new features like YAML configuration format support, listing of available updates and security fixes. The newly created tool should eventually replace the previous ones. Natasha’s mentors: Kristin Whetstone, Jon Brohauge and Arnab Banerjee. Also, many contributors from Platform SIG and JCasC plugin team joined the project as a key stakeholders and subject-matter experts.

Plugin Manager Tool YAML file

Working Hours Plugin - UI Improvements

Jenkins UI and frontend framework are a common topic in the Jenkins project, especially in recent months after the new UX SIG was established. Jack Shen was working on exploring new ways to build Jenkins Web UI together with his mentor Jeff Pearce. Jack updated the Working Hours Plugin to use UI controls provided by standard React libraries. Then he documented his experienced and a created template for plugins with React-based UI.

Web UI controls in React

Remoting over Apache Kafka with Kubernetes features

Long Le Vu Nguyen was working on extended Kubernetes support in the Remoting over Apache Kafka Plugin. His mentors were Andrey Falco and Pham vu Tuan who was our GSoC 2018 student and the plugin creator. During this project Long has added a new agent launcher which provisions Jenkins agents in Kubernetes and connects them to the master. He also created a Cloud API implementation for it and a new Helm chart which can provision Jenkins as entire system in Kubernetes, with Apache Kafka enabled by default. All these features were released in Remoting over Apache Kafka Plugin 2.0.

Jenkins in Kubernetes with Apache Kafka

Multi-branch Pipeline support for Gitlab SCM

Parichay Barpanda was working on the new GitLab Branch Source Plugin with Multi-branch Pipeline Jobs and Folder Organisation support. His mentors were Marky Jackson-Taulia, Justin Harringa, Zhao Xiaojie and Joseph Petersen. The plugin scans the projects, importing the pipeline jobs it identifies based on the criteria provided. After a project is imported, Jenkins immediately runs the jobs based on the Jenkinsfile pipeline script and notifies the status to GitLab Pipeline Status. This plugin also provides GitLab server configuration which can be configured in Configure System or via Jenkins Configuration as Code (JCasC). read more about this project in the GitLab Branch Source 1.0 announcement.

Gitlab Multi-branch Pipeline support

Projects which were not completed

Not all projects have been completed this year. We were also working on Artifact Promotion plugin for Jenkins Pipeline and on Cloud Features for External Workspace Manager Plugin, but unfortunately both projects were stopped after coding phase 1. Anyway we got a lot of experience and takeaways in these areas (see linked Jira tickets!), and we hope that these stories will be implemented by Jenkins contributors at some point. Google Summer of Code 2020 maybe?

Running the GSoC program at our organization level

Here are some of the things our organization did before and during GSoC behind the scenes. To prepare for the influx of students, we updated all our GSoC pages and wrote down all the knowledge we accumulated over the years of running the program. We started preparing in October 2018, long before the official start of the program. The main objective was to address the feedback we got during GSoC 2018 retrospectives.

Project ideas. We started gathering project ideas in the last months of 2018. We prepared a list of project ideas in a Google doc, and we tracked ownership of each project in a table of that document. Each project idea was further elaborated in its own Google doc. We find that when projects get complicated during the definition phase, perhaps they are really too complicated and should not be done.

Since we wanted all the project ideas to be documented the same way, we created a template to guide the contributors. Most of the project idea documents were written by org admins or mentors, but occasionally a student proposed a genuine idea. We also captured contact information in that document such as GitHub and Gitter handles, and a preliminary list of potential mentors for the project. We embedded all the project documents on our website.

Mentor and student guidelines. We updated the mentor information page with details on what we expect mentors to do during the program, including the number of hours that are expected from mentors, and we even have a section on preventing conflict of interest. When we recruit mentors, we point them to the mentor information page.

We also updated the student information page. We find this is a huge time saver as every student contacting us has the same questions about joining and participating in the program. Instead of re-explaining the program each time, we send them a link to those pages.

Application phase. Students started to reach out very early on as well, many weeks before GSoC officially started. This was very motivating. Some students even started to work on project ideas before the official start of the program.

Project selection. This year the org admin team had some very difficult decisions to make. With lots of students, lots of projects and lots of mentors, we had to request the right number of slots and try to match the projects with the most chances of success. We were trying to form mentor teams at the same time as we were requesting the number of slots, and it was hard to get responses from all mentors in time for the deadline. Finally we requested fewer slots than we could have filled. When we request slots, we submit two numbers: a minimum and a maximum. The GSoC guide states that:

  • The minimum is based on the projects that are so amazing they really want to see these projects occur over the summer,

  • and the maximum number should be the number of solid and amazing projects they wish to mentor over the summer.

We were awarded minimum. So we had to make very hard decisions: we had to decide between "amazing" and "solid" proposals. For some proposals, the very outstanding ones, it’s easy. But for the others, it’s hard. We know we cannot make the perfect decision, and by experience, we know that some students or some mentors will not be able to complete the program due to uncontrollable life events, even for the outstanding proposals. So we have to make the best decision knowing that some of our choices won’t complete the program.

Community Bonding. We have found that the community bonding phase was crucial to the success of each project. Usually projects that don’t do well during community bonding have difficulties later on. In order to get students involved into the community better, almost all projects were handled under umbrella of Special Interest Groups so that there were more stakeholders and communications.

Communications. Every year we have students who contact mentors via personal messages. Students, if you are reading this, please do NOT send us personal messages about the projects, you will not receive any preferential treatment. Obviously, in open source we want all discussions to be public, so students have to be reminded of that regularly. In 2019 we are using Gitter chat for most communications, but from an admin point of view this is more fragmented than mailing lists. It is also harder to search. Chat rooms are very convenient because they are focused, but from an admin point of view, the lack of threads in Gitter makes it hard to get an overview. Gitter threads were added recently (Nov 2019) but do not yet work well on Android and iOS. We adopted Zoom Meetings towards the end of the program and we are finding it easier to work with than Google Hangouts.

Status tracking. Another thing that was hard was to get an overview of how all the projects were doing once they were running. We made extensive use of Google sheets to track lists of projects and participants during the program to rank projects and to track statuses of project phases (community bonding, coding, etc.). It is a challenge to keep these sheets up to date, as each project involves several people and several links. We have found it time consuming and a bit hard to keep these sheets up to date, accurate and complete, specially up until the start of the coding phase.

Perhaps some kind of objective tracking tool would help. We used Jenkins Jira for tracking projects, with each phase representing a separate sprint. It helped a lot for successful projects. In our organization, we try to get everyone to beat the deadlines by a couple of days, because we know that there might be events such as power outages, bad weather (happens even in Seattle!), or other uncontrolled interruptions, that might interfere with submitting project data. We also know that when deadlines coincide with weekends, there is a risk that people may forget.

Retrospective. At the end of our project, we also held a retrospective and captured some ideas for the future. You can fine the notes here. We already addressed the most important comments in our documentation and project ideas for the next year.

Recognition

Last year, we wanted to thank everyone who participated in the program by sending swag. This year, we collected all the mailing addresses we could and sent to everyone we could the 15-year Jenkins special edition T-shirt, and some stickers. This was a great feel good moment. I want to personally thank Alyssa Tong her help on setting aside the t-shirt and stickers.

swag before shipping

Mentor summit

Each year Google invites two or more mentors from each organization to the Google Summer of Code Mentor Summit. At this event hundreds of open-source project maintainers and mentors meet together and have unconference sessions targeting GSoC, community management and various tools. This year the summit was held in Munich, and we sent Marky Jackson and Oleg Nenashev as representatives there.

Apart from discussing projects and sharing chocolate, we also presented Jenkins there, conducted a lightning talk and hosted the unconference session about automation bots for GitHub. We did not make a team photo there, so try to find Oleg and Marky on this photo:

GSoC2019 Mentor summit

GSoC Team at DevOps World | Jenkins World

We traditionally use GSoC organization payments and travel grants to sponsor student trips to major Jenkins-related events. This year four students traveled to the DevOps World | Jenkins World conferences in San-Francisco and Lisbon. Students presented their projects at the community booth and at the contributor summits, and their presentations got a lot of traction in the community!

Thanks a lot to Google and CloudBees who made these trips possible. You can find a travel report from Natasha Stopa here, more travel reports are coming soon.

gsoc2019 team jw us gsoc2019 team jw lisbon

Conclusion

This year, five projects were successfully completed. We find this to be normal and in line with what we hear from other participating organizations.

Taking the time early to update our GSoC pages saved us a lot of time later because we did not have to repeat all the information every time someone contacted us. We find that keeping track of all the mentors, the students, the projects, and the meta information is a necessary but time consuming task. We wish we had a tool to help us do that. Coordinating meetings and reminding participants of what needs to be accomplished for deadlines is part of the cheerleading aspect of GSoC, we need to keep doing this.

Lastly, I want to thank again all participants, we could not do this without you. Each year we are impressed by the students who do great work and bring great contributions to the Jenkins community.

GSoC 2020?

Yes, there will be Google Summer of Code 2020! We plan to participate, and we are looking for project ideas, mentors and students. Jenkins GSoC pages have been already updated towards the next year, and we invite everybody interested to join us next year!

Wednesday, 29 January 2020

Five Ways Big Data is Invaluable to Launching an Online Business

Big data is incredibly valuable for businesses. The market is expected to reach $274 billion in the next two years. Online businesses in particular stand to benefit from big data, due to the logistics of their business models.Every entrepreneur has some sort of online presence these days since most of us are on social media or have our own website. The good news is that big data is making it easier to create a successful business model.It’s become so easy for people to host and share their ideas online that having your own website is now actually pretty common. It doesn’t have to be the full-time gig it used to be - a lot of people run them alongside their day jobs, whether it’s for a hobby or even a means to make an easy side hustle. You can make things even easier by using big data to automate many aspects of your business.Anyone can make a website but if you’re looking to make it profitable or make it the basis for your online business, you’ll want to take it more seriously than just a project to chip away at in your spare time. Running a successful online business requires ...


Read More on Datafloq

Query Delta Lake Tables from Presto and Athena, Improved Operations Concurrency, and Merge performance

We are excited to announce the release of Delta Lake 0.5.0, which introduces Presto/Athena support and improved concurrency.

The key features in this release are:

  • Support for other processing engines using manifest files (#76) – You can now query Delta tables from Presto and Amazon Athena using manifest files, which you can generate using Scala, Java, Python, and SQL APIs. See the Presto and Athena to Delta Lake Integration documentation for details
  • Improved concurrency for all Delta Lake operations (#9, #72, #228) – You can now run more Delta Lake operations concurrently. Delta Lake’s optimistic concurrency control has been improved by making conflict detection more fine-grained. This makes it easier to run complex workflows on Delta tables. For example:
    • Running deletes (e.g. for GDPR compliance) concurrently on older partitions while newer partitions are being appended.
    • Running updates and merges concurrently on disjoint sets of partitions.
    • Running file compactions concurrently with appends (see below).

For more information, please refer to the open-source Delta Lake 0.5.0 release notes. In this blog post, we will elaborate on reading Delta Lake tables with Presto, improved operations concurrency, easier and faster data deduplication using insert-only merge.

Reading Delta Lake Tables with Presto

As described in Simple, Reliable Upserts and Deletes on Delta Lake Tables using Python APIs, modifications to the data such as deletes are performed by selectively writing new versions of the files containing the data be deleted and only marks the previous files as deleted. The advantage of this approach is that Delta Lake enables us to travel back in time (i.e. time travel) and query previous versions.

To understand which files (and rows) contain the latest data, by default you can query the transaction log (more information at Diving Into Delta Lake: Unpacking The Transaction Log). Other systems like Presto and Athena can read a generated manifest file – a text file containing the list of data files to read for querying a table. To do this, we will follow the Python instructions; for more information, refer to Set up the Presto or Athena to Delta Lake integration and query Delta tables.

Generate Delta Lake Manifest File

Let’s start by creating the Delta Lake manifest file with the following code snippet.


deltaTable = DeltaTable.forPath(pathToDeltaTable)
deltaTable.generate("symlink_format_manifest")

As the name implies, this generates the manifest file in the table root folder. If you had created the departureDelays table per Simple, Reliable Upserts and Deletes on Delta Lake Tables using Python APIs, you will have a new folder in the table root folder:


$/departureDelays.delta/_symlink_format_manifest

with a single file named manifest. If you review the files within the manifest (e.g. cat manifest), you will get the following output indicating the files that contain the latest snapshot.


file:$/departureDelays.delta/part-00003-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00006-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00001-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00000-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00000-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00001-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00002-...-c000.snappy.parquet
file:$/departureDelays.delta/part-00007-...-c000.snappy.parquet

Create Presto Table to Read Generated Manifest File

The next step is to create an external table in the Hive Metastore so that Presto (or Athena with Glue) can read the generated manifest file to identify which Parquet files to read for reading the latest snapshot of the Delta table. Note, for Presto, you can either use Apache Spark or the Hive CLI to run the following command. k.


1. CREATE EXTERNAL TABLE departureDelaysExternal ( ... )
2. ROW FORMAT SERDE
   'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe'
3. STORED AS INPUTFORMAT
4. OUTPUTFORMAT
   'org.apache.hadoop.hive.ql.io.HiveIgnoreKeyTextOutputFormat'
5. LOCATION '$/departureDelays.delta/_symlink_format_manifest'

Some important notes on schema enforcement:

  • The schema defined on line 1 must match the schema of the Delta Lake table (e.g. in this example, departureDelaysExternal). Note, the partitioning scheme is optional.
  • Line 5 points to the location of the manifest file in the form of /_symlink_format_manifest/

The SymlinkTextInputFormat configures Presto (or Athena) to get the list of Parquet data files from the manifest file instead of using directory listing. Note, for partitioned tables, there are additional steps that will need to be performed per Configure Presto to read the generated manifests.

Update the Manifest File

It is important to note that every time the data is updated, you will need to regenerate the manifest file so Presto will be able to see the latest data.

Improved Operations Concurrency

With the following pull requests, you can now run even more Delta Lake operations concurrently. With finer grain conflict detection, these updates make it easier to run complex workflows on Delta tables such as:

  • Running deletes (e.g. for GDPR compliance) concurrently on older partitions while newer partitions are being appended.
  • Running file compactions concurrently with appends.
  • Running updates and merges concurrently on disjoint sets of partitions.

Concurrent Appends Use Cases

For example, typically there is a ConcurrentAppendException thrown during concurrent merge operations when concurrent transaction adds records to the same partition.


// Target 'deltaTable' is partitioned by date and country
deltaTable.as("t").merge(
    source.as("s"),
    "s.user_id = t.user_id AND s.date = t.date AND s.country = t.country")
  .whenMatched().updateAll()
  .whenNotMatched().insertAll()
  .execute()

The above code snippet potentially can cause conflicts because the condition is not explicit enough resulting even though the table is already partitioned by date and country. The issue is that the query currently will scan the entire table potentially resulting in a conflict with concurrent operations updating any other partitions. By specifying specificDate and specificCountry so you can merge on a specific date or country, this operation is now safe to run concurrently on different dates and countries.


// Target 'deltaTable' is partitioned by date and country
deltaTable.as("t").merge(
    source.as("s"),
    "s.user_id = t.user_id AND d.date = '" + specificDate + "' AND d.country = '" + specificCountry + "'")
  .whenMatched().updateAll()
  .whenNotMatched().insertAll()
  .execute()

This approach is the same for all other Delta Lake operations (e.g. delete, metadata changed, etc.).

Concurrent File Compaction

If you are continuously writing data to a Delta table, over time a large number of files will be accumulated. This is especially important in streaming scenarios as you are adding data in small batches. This results in the file system continuing to accumulate many small files; this will degrade query performance over time. An important optimization task is to periodically take a large number of small files and rewrite them to a smaller number of larger files, i.e. file compaction.

In the past, there was a higher potential for an exception when concurrently querying the data and running file compaction. But, because of these improvements, you can also run queries (including streaming queries) and file compaction concurrently without any exceptions. For example, If your table is partitioned and you want to repartition just one partition based on a predicate, you can read only the partition using where and write back to that using replaceWhere:


path = "..."
partition = "year = '2019'"
numFilesPerPartition = 16   # Compact partition of a table to no. of files

(spark.read
  .format("delta")
  .load(path)
  .where(partition)
  .repartition(numFilesPerPartition)
  .write
  .option("dataChange", "false")
  .format("delta")
  .mode("overwrite")
  .option("replaceWhere", partition)
  .save(path))

Note, use the dataChange == false option only when there are no data changes (such as in the preceding code snippet) otherwise this may corrupt the underlying data.

Easier and Faster Data Deduplication Using Insert-only Merge

A common ETL use case is to collect logs and append them into a Delta Lake table. A common issue is that the source generates duplicate log records. With Delta Lake merge, you can avoid inserting these duplicate records such as the following code snippet involving merging updated flight data.


# Merge merge_table with flights
deltaTable.alias("flights") \
    .merge(merge_table.alias("updates"),"flights.date = updates.date") \
    .whenMatchedUpdate(set = { "delay" : "updates.delay" } ) \
    .whenNotMatchedInsertAll() \
    .execute()

Prior to Delta Lake 0.5.0, it was not possible to read deduped data as a stream from a Delta Lake table because insert-only merges were not pure appends into the table.

For example, in a streaming query, you can run a merge operation in foreachBatch to continuously write any streaming data into a Delta Lake table with deduplication as noted in the following PySpark snippet.


from delta.tables import *

deltaTable = DeltaTable.forPath(spark, "/data/aggregates")

# Function to upsert microBatchOutputDF into Delta table using merge
def upsertToDelta(microBatchOutputDF, batchId):
  deltaTable.alias("t").merge(
      microBatchOutputDF.alias("s"),
      "s.key = t.key") \
    .whenMatchedUpdateAll() \
    .whenNotMatchedInsertAll() \
    .execute()
}

# Write the output of a streaming aggregation query into Delta table
streamingAggregatesDF.writeStream \
  .format("delta") \
  .foreachBatch(upsertToDelta) \
  .outputMode("update") \
  .start()

In another streaming query, you can continuously read deduplicated data from this Delta Lake table. This is possible because insert-only merge – introduced in Delta Lake 0.5.0 – will only append new data to the Delta table.

Getting Started with Delta Lake 0.5.0

Try out Delta Lake today by trying out the preceding code snippets on your Apache Spark 2.4.3 (or newer) instance. By using Delta Lake, you can make your data lakes more reliable (whether you create a new one or migrate an existing data lake). To learn more, refer to https://delta.io/ and join the Delta Lake open source community via Slack and Google Group. You can track all the upcoming releases and planned features in Delta Lake github milestones.

Credits

We want to thank the following contributors for updates, doc changes, and contributions in Delta Lake 0.5.0: Andreas Neumann, Andrew Fogarty, Burak Yavuz, Denny Lee, Fabio B. Silva, JassAbidi, Matthew Powers, Mukul Murthy, Nicolas Paris, Pranav Anand, Rahul Mahadev, Reynold Xin, Shixiong Zhu, Tathagata Das, Tomas Bartalos, and Xiao Li.

--

Try Databricks for free. Get started today.

The post Query Delta Lake Tables from Presto and Athena, Improved Operations Concurrency, and Merge performance appeared first on Databricks.

Data Lineage - How the History of Your Data can Influence the Future of Your Business

Almost every single industry has been revolutionized to at least some degree by data processing. Whether you manage a local restaurant or a massive distributorship, there's a good chance that you've recently made several decisions based on the availability of data that you previously lacked.Fewer people are paying attention to where their data is coming from, however. In fact, business owners who rely on databases are likely to accept everything they read as the truth. Unfortunately, they don't realize that there are certain coding tolerances involved with collecting information on any digital device.In some cases, an insecure system can actually suffer from bogus information as well. That makes data lineage extremely important to ensure that you are getting the most accurate numbers possible.Implementing a Data Lineage PlanQuite a few people are probably asking the question "what is Data Lineage" because, at the moment, few organizations are investing much into this field. That's quite a shame considering how much promise it shows for those who need to check the quality of the numbers they're getting.Data lineage is essentially the life cycle that every piece of digital information you collect goes through. This includes where the data comes from, how it got ...


Read More on Datafloq

Data Science: Breaking Down The Data Silos

In today’s digitized economy, the capability to utilize data represents a real and indispensable competitive advantage. Organizations are using advanced technologies to unbolt the true value of their data. This allows them to make smart decisions, accelerate innovation, serve up better customer experiences, and respond to problems faster. The main issue standing in the way of transforming the deluge of big data into decision-making is- Data Silos. Presently, every department has its data repository which is isolated from the rest of the organization, resulting in - many conflicting versions of the truth. Silos have made it difficult to analyze, manage, & activate data. So, here we are to discuss how you can break down the Silos and activate the power of Business.What are Data Silos?Data silo refers to independent sets of data within an organization. Often aligned to either IT systems or business functions, data silos are where only a limited group of people have knowledge or access to the data resources available. It is a big problem for organizations that are looking out for sheer productivity, business agility, and efficiency. Breaking down your data silos starts with knowing how they were initially created? Eliminating data silos can help you ...


Read More on Datafloq

Analytics and Artificial Intelligence – A Blue Ocean of Opportunities

“A whole new world …” the lyrics go as Aladdin flies up in the skies on his magical carpet. Yes, a childhood nostalgia, but yes today it can become a business reality. Yes, of course, with the winning combination of Analytics and Artificial Intelligence (AI).Any business today, you name it, encounters a highly competitive landscape. However, the data that it generates, both structured and unstructured, has a kind of a digital footprint, which can be thoughtfully analyzed and utilized. Though, you need that vantage point to correlate the dots and create the bigger picture. It is the same bird’s eye view, which the above lyrics hint at and allows you to join the dots to identify those blue ocean opportunities.How does this technology combination work?Businesses typically have their own enterprise data management platform to manage, store, and retrieve their multi-structured data. They can use Analytics and AI on top of it to correlate their data sets and come up with evolving business trends and patterns. This AI-powered Analytics offers a practical business solution to what otherwise takes months of tedious, manual, and time-consuming analysis. This ...


Read More on Datafloq

How AI and IoT will revolutionize shopping in the future?

Let me take you 20 years back when online shopping just kicked off and started to gain popularity. And now, 20 years later it is at the boom when everything can be bought in just one click from anywhere. Not only that, computers have become powerful plus cheaper and the internet faster. I am not only talking about metropolitan areas! Technology and its roots have spread everywhere. Online shopping might rock at this particular time but digital shopping is something which is rising slowly too. Artificial intelligence (AI), augmented reality (AR), virtual reality (VR), and the Internet of Things (IoT) are all the driving forces to revolutionize slash change the way we shop.The Internet of ThingsThis is not as scary as the term sounds. Instead, we all are using it and getting advantages from this keyword, which is called the 'Internet of Things.' From Amazon’s Echo to Google Assistant, you have used these devices to play a song early morning, getting news or look up the weather. These home assistants have made our life so easier. You don't need to turn on the TV, go to your phone for playing music - just give a command and you will see ...


Read More on Datafloq

Bengaluru: A Gateway of South Indian Delicacies

Bengaluru is well known for its South Indian delicacies, however, with an emergence and demand for North Indian delicacies have kept the city to delight with taste. Bengaluru is an inconsistency on each turn. It is as yet reminiscent of the tired old town which was out of the blue tossed into the bewildering universe of a clamoring metropolitan. Old Bangloreans frequently regret the populace that came pouring in with the city turning into an IT center point, and all the more as of late, a startup center point. Traffic turned into a bad dream in a city where the sentimental rear ways couldn't deal with the swelling weights of a huge number of extra vehicles. Be that as it may, with any change, there additionally come its very own arrangement of points of interest which you can also explore in your city by ordering with Ovenstory Coupons. Here are few things that we recommend you attempt when in the city: Local Meals at Vidhyarthi Bhavan This is an old most loved and one you should accomplish for sentimentality. Not a spot for listless dinners, you share tables, get served in a jiffy and get moving. The masala dosa is ...


Read More on Datafloq

How Artificial Intelligence in Enhancing Digital KYC?

Artificial Intelligence has many innovative solutions that cover different sectors and transform facilities like Know Your Customer (KYC) and Anti-money Laundering (AML) standards. AI systems are managed which help in tracking the malicious transactions and associated present and future risks based on the overall monitoring perspectives which ultimately help in combating risks of fraud. To know what is cooking inside the system an intelligent decision making is required that comes up with observations that seem dangling from the legitimate path.With AI technology the identification of high-risks has become easy. To perform multitasking that can be replaced with human efforts that are utilized for manual processing by AI. AI not only saves them money but efforts and valuable time of employees can be spared to put in some other productive organizational tasks. The knowledge of Natural Language Processing (NLP) and Machine Learning (ML) is being used by AI so this hybrid technology enhances the operations efficiently. Particularly in labor-intensive areas, this fusion results in an efficient automated system. The traditional institutions which include financial institutions, banks, insurance companies, etc. are in dire need of complying with KYC/AML procedures that can take advantage of these growing technologies. Understanding the result accuracy ...


Read More on Datafloq

Tatas look to build EV ecosystem

The ecosystem, appropriately named Tata UniEVerse, comprises participation from multiple Tata Group companies, including Tata Power, Tata Chemicals, Tata AutoComp, Tata Motors Finance, Tata Consultancy Services, and Croma.

Isro to launch 10 satellites to replace its ageing fleet

The plan includes launching high throughput satellites that can beam high-speed internet at more than 300 gigabytes per second into remote corners

Isro satellite data to be used for government planning services

Isro's geospatial data will help the government better plan e-government services in villages.

Tuesday, 28 January 2020

On-Demand Webinar: Geospatial Analytics and AI in the Public Sector

We recently hosted a live webinar — Geospatial Analytics and AI in Public Sector — during which we covered top geospatial analysis use cases in the Public Sector along with live demos showcasing how to build scalable analytics and machine learning pipelines on geospatial data at sale.

Geospatial Analytics and AI in Public Sector Webcast

Geospatial Analytics Webinar Overview

Today, government agencies have access to massive volumes of geospatial information that can be analyzed to deliver on a broad range of decision-making and predictive analytics use cases from transportation planning to disaster recovery and population health management.

While many agencies have invested in geographic information systems that produce volumes of geospatial data, few have the proper technology and technical expertise to prepare these large, complex datasets for analytics — inhibiting their ability to build AI applications.

In this webinar, we reviewed:

  • Top geospatial big data use cases in Public Sector spanning public safety, defense, infrastructure management, health services, fraud prevention and more
  • Challenges analyzing large volumes of geospatial data with legacy architectures
  • How Databricks and open-source tools can be used to overcome these challenges in the cloud
  • Technical demos and notebooks shared on the webinar:
    • Object Detection in xView Imagery: Bridges complex object detection using Deep Learning with accessible SQL-based analytics for non-data scientist personas. Download related notebooks: data engineering and analysis.
    • Processing Large-Scale NYC Taxi Pickup / Dropoff Vectors: Optimizes geospatial predicate operations and joins to associate raw pick-up/drop-off coordinates with their corresponding NYC neighborhood boundaries to facilitate spatial analysis. Download related notebook.

If you’d like free access to the Unified Data Analytics Platform and try our notebooks on it, you can access a free Databricks trial here.

At the end of the webinar we held a Q&A. Below are the questions and answers:

Q: We deal with large volumes of streaming geospatial data. How would you recommend handling these real-time data streams for downstream analytics?

A: This can be broken down to (1) handling large volumes of streaming data and (2) performing downstream geospatial analytics. Databricks makes processing and storing large volumes of streaming data simple, reliable, and performant. Please reference Delta Lake on Databricks and Introduction to Delta Lake for some additional material. The second part builds on storage and schema decisions made during the processing phase. Spatial analysis is fundamentally addressed through the use of Spark SQL, DataFrames, and Datasets to power transformations and actions over data originating from various formats and schemas. Databricks offers various runtimes such as Machine Learning Runtime and Databricks Runtime with Conda which pre-bundle popular libraries including Tensorflow, Horovod, PyTorch, Scikit-Learn, and Anaconda for both CPU and GPU clusters to facilitate common Data Engineering and Data Science needs. Customers can also manage their own Libraries or Containers to customize the environment for any analytic, to include spatial specific needs. Please reference popular spatial frameworks listed in the following question as well as the FINRA Customer Case Study

Q: What are some of the more popular spatial frameworks being used in the public sector?

A: Popular frameworks which extend Apache Spark for geospatial analytics include GeoMesa, GeoTrellis, Rasterframes, and GeoSpark. In addition, Databricks makes it easy to use single-node libraries such as GeoPandas, Shapely, Geospatial Data Abstraction Library (GDAL), and Java Topology Service (JTS). By wrapping function calls in user-defined functions (UDFs) these libraries can further be leveraged in a distributed context as well. UDFs offer a simple approach for scaling existing workloads with minimal code changes.

Q: Where is my data stored and how does Databricks help ensure data security?

A: Your data is stored in your own cloud data lake, such as in AWS S3 or Azure Blob Storage. However, data lakes often have data quality issues, due to a lack of control over ingested data. Delta Lake adds a storage layer to data lakes to manage data quality, ensuring data lakes contain only high-quality data for consumers. Delta Lake also offers capabilities like ACID transactions to ensure data integrity with serializability as well as audit history, allowing you to maintain log records details about every change made to data, providing a full history of changes, for compliance, audit, and reproduction. Additionally, Delta Lake has been designed to address various right-to-erasure initiatives such as the General Data Protection Regulation (GDPR) and recently the California Consumer Privacy Act (CCPA), reference Make Your Data Lake CCPA Compliant with a Unified Approach to Data and Analytics. As part of our Enterprise Cloud Service, Delta Lake is tightly integrated with other Databricks Enterprise Security features.

Additional Geospatial Analytics Resources

--

Try Databricks for free. Get started today.

The post On-Demand Webinar: Geospatial Analytics and AI in the Public Sector appeared first on Databricks.

Machine Learning Can Assist with Five Year Balance Sheet Forecasts

Machine learning is making many business processes much easier. A growing number of companies are using machine learning to create more accurate financial documents by leveraging sophisticated machine learning technology.The most astute entrepreneurs are using machine learning to predict the future equity on their balance sheets. This can be very useful for business owners that intend to solar companies over the next few years. While I was working on my MBA, the majority of my classmates agreed that their most hated class was financial management. The average person does not have the aptitude for math needed to create complicated financial documents. Even the mathematically gifted people that choose the financial track of the business world tend to get burnt out performing extensive calculations. The financial forecasting process is one of the least liked aspects of business. Business professionals that don’t enjoy this process are compelled to use technology to minimize the challenges that it entails. The benefits of machine learning in the financial world have been well documented for over 20 years. One study from 1998 showed that machine learning tools can predict an increase in the stock market with a 93.3% accuracy. During the same year, another team of ...


Read More on Datafloq

Object Transplants Can Confound AI Machine Learning And Autonomous Cars

By Lance Eliot, the AI Trends Insider

Have you ever seen the videos that depict a scene in which there is some kind of activity going on such as people tossing a ball to each other and then a gorilla saunters through the background?

If you’ve not seen any such video, it either means you are perhaps watching too many cat videos or that you’ve not yet been introduced to the concept of inattentional blindness or what some consider an example of selective attention.

The gorilla is not a real gorilla, but instead a person in a gorilla suit (wanted to mention this in case you were worried that an actual gorilla was invading human gatherings and that the planet might be headed to the apes!).

Overall, the notion is that you become focused on the other activity depicted in the video and fail to notice that a gorilla has ambled into the scene.

Gorilla On The Mind

When I tell people about this phenomenon, and if they’ve not yet seen one of these videos, they are often quite doubtful that those viewing such videos really did not see the gorilla.

Of course, now that I’ve told you about it, you are not likely to be “fooled” by any such video since you have been forewarned about it.

Sorry about that.

Guess I should have said spoiler alert before I told you about the gorilla.

In any case, for those doubters about the assertion that people watching such a video are apt to not notice a gorilla (which, not noticing a squirrel might seem more plausible, I acknowledge), I assure you that a large number of cognition related experiments have been done with the “invisible gorilla” videos and such studies support this claim. The experiments tend to indicate that many people watching the video are completely unaware that a gorilla moseyed into the scene.

In fact, if you ask such people immediately after viewing the video whether they noticed a gorilla, most people swear it is outright impossible that a gorilla was in the scene.

Upon showing them the video a second time, they then see the gorilla, but also insist that you are tricking them by showing an identical video but one that you’ve sneakily inserted a gorilla into after-the-fact. These people will be utterly convinced that you are trying to trick them into falsely believing that the original video had a gorilla in it.

One experiment even involved videoing the person as they initially watched the gorilla related video, so that after being told about the gorilla, the person could watch the video of themselves watching the video, and hopefully then believe that there was a gorilla in the initial video.

I’m sure that there are some so suspicious that they figure you doctored the video of the video and thus refuse to still believe what they saw (or, what they didn’t see).

There is a bit of trickery somewhat involved in this matter.

Most of the time, the experimenter tells the person to be watching for something else in the video, such as watching to see if anyone drops the ball while tossing it or counting how many times the ball is tossed back-and-forth. I mention this rather important point because if you had no other focus related to the video, the odds are much higher that you would notice the gorilla. Because your attention has been purposely shaped though by the experimenter, you tend to block out other aspects that don’t pertain to the matter you were directed to pay attention to.

From a cognitive perspective, the aspect of not noticing the gorilla is at times attributed to inattentional blindness.

This means that you were not particularly paying attention to spot the gorilla and were therefore blind to noticing it. Others tend to describe this as an example of selective attention. Your attention is focused on something else that you were instructed to watch for. As a result, you select just those aspects in the scene that are related to the needed focus. No need to watch for other aspects and in fact it could be distracting if you did look elsewhere in the scene and therefore you might end-up doing worse on the assigned task at-hand.

The Elephant In The Room

Let’s switch now from talking about gorillas to instead talking about elephants.

You’ve likely heard the famous expression about there being an elephant in the room.

This is a popular metaphorical idiom and means that there is something that everyone notices or is aware of, but for which no one wants to bring it up or talk about it.

For example, I went to an evening party the other night and one of the attendees arrived wearing a quite unusual hat. The attendees were all exceedingly reserved and civil, and no one overtly pointed at the hat or made any direct explicit remarks about the hat. The hat was the elephant in the room. It was there but no one particularly spoke of it. We all saw it and knew that it was unusual.

This summer, I went to our local county fair and was able to stand in a room with an elephant. Yes, an actual elephant. In this case, everyone knew there was an elephant in the room and everyone spoke up about it and pointed at it. This was not a metaphorical elephant. Though, it certainly would have been more interesting if the elephant had been in the room and everyone pretended to not notice. I guess that might be somewhat dangerous though, if the elephant suddenly decided to lumber around in the room.

Anyway, some AI researchers recently conducted a fascinating study about an elephant in the room (of sorts).

It was a picture of an elephant.

The experiment they conducted dealt with the use of Machine Learning (ML) and Artificial Neural Networks (ANN).

Allow me to elaborate.

First, you might be aware that when using Machine Learning and deep neural networks that you need to contend with so-called adversarial examples.

This refers to the notion that some ML and ANN’s can be hyper-sensitive to targeted perturbations.

Let’s suppose you opt to train a neural network to be able to find images of turtles.

Little turtles, big turtles, turtles hiding in their shells, turtles poking their heads out, etc. You start by showing the neural network hundreds or maybe thousands of relatively crisp and clear-cut pictures of turtles. The neural network seems to gradually be getting pretty good at picking out the turtle in the pictures being shown to it. You test this by showing the neural network a picture that has a turtle and a squirrel, and sure enough the neural network is able to spot which of those animals is the turtle (and doesn’t misidentify the squirrel as being a turtle).

Great, you are ready to use the neural network.

But, suppose I do a bit of Photoshop work on the picture that contained the turtle and the squirrel. I make a copy of the squirrel’s tail and paste it onto the posterior of the turtle. I copy one of the legs of the turtle onto the squirrel. Admittedly, this looks somewhat Frankenstein-like, but go with me on this notion for the moment.

I show the picture to the neural network.

What could happen is that the neural network now reports that the turtle is not a turtle and asserts that there is not a turtle at all in the picture. Or, the neural network could assert that there are two turtles in the picture, falsely believing that the squirrel is also a turtle. How could this happen?

Sometimes, even a small perturbation can confuse the neural network.

If the neural network was focusing on only the legs of turtles, and it detected a turtle leg on the squirrel, it might then conclude that the squirrel is also a turtle. If the neural network had been honing on the tail of the turtle to determine whether a turtle is a turtle, the aspect that the turtle had a squirrel’s tail would have caused the neural network to no longer believe that the turtle is a turtle.

This is one of the known dangers, or let’s say inherent limitations, about the use of machine learning and neural networks.

A good AI developer will try to ascertain the sensitivity of the neural network to the various “factors” that the neural network had landed upon to do its detective work. Unfortunately, many deep neural networks are so complicated that it is not readily feasible to determine what it is using to do the detective work. All you have is a complicated set of mathematical aspects and for which you might not be able to “logically” discern what it refers to.

I’ve mentioned many times that this is something that AI self-driving cars need to be attuned to.

AI Autonomous Cars And Inattentive Attention

There are various studies that have shown how easy it can be to confuse a machine learning algorithm or neural network that might be used by an AI self-driving driverless autonomous car.

These are often used when the AI is examining the visual camera images being collected by the sensors of the self-driving car.

For several of my articles about notable machine learning and neural network limitations and concerns related to AI self-driving cars, see these:

https://aitrends.com/selfdrivingcars/expert-systems-ai-self-driving-cars-crucial-innovative-techniques/

https://aitrends.com/selfdrivingcars/ensemble-machine-learning-for-ai-self-driving-cars/

https://aitrends.com/ai-insider/machine-learning-benchmarks-and-ai-self-driving-cars/

https://aitrends.com/ai-insider/explanation-ai-machine-learning-for-ai-self-driving-cars/

https://aitrends.com/ai-insider/federated-machine-learning-for-ai-self-driving-cars/

At the Cybernetic AI Self-Driving Car Institute, we are developing AI software for self-driving cars. As such, we are actively using and developing these AI systems via the use of machine learning and neural networks. It’s important for the auto makers and tech firms to be using such tools carefully and wisely.

Allow me to elaborate.

First, I’d like to clarify and introduce the notion that there are varying levels of AI self-driving cars. The topmost level is considered Level 5. A Level 5 self-driving car is one that is being driven by the AI and there is no human driver involved. For the design of Level 5 self-driving cars, the automakers are even removing the gas pedal, the brake pedal, and steering wheel, since those are contraptions used by human drivers. The Level 5 self-driving car is not being driven by a human and nor is there an expectation that a human driver will be present in the self-driving car. It’s all on the shoulders of the AI to drive the car.

For self-driving cars less than a Level 5 and Level 4, there must be a human driver present in the car. The human driver is currently considered the responsible party for the acts of the car. The AI and the human driver are co-sharing the driving task. In spite of this co-sharing, the human is supposed to remain fully immersed into the driving task and be ready at all times to perform the driving task. I’ve repeatedly warned about the dangers of this co-sharing arrangement and predicted it will produce many untoward results.

For my overall framework about AI self-driving cars, see my article: https://aitrends.com/selfdrivingcars/framework-ai-self-driving-driverless-cars-big-picture/

For the levels of self-driving cars, see my article: https://aitrends.com/selfdrivingcars/richter-scale-levels-self-driving-cars/

For why AI Level 5 self-driving cars are like a moonshot, see my article: https://aitrends.com/selfdrivingcars/self-driving-car-mother-ai-projects-moonshot/

For the dangers of co-sharing the driving task, see my article: https://aitrends.com/selfdrivingcars/human-back-up-drivers-for-ai-self-driving-cars/

Let’s focus herein on the true self-driving car.

Much of the comments apply to the less than Level 5 and Level 4 self-driving cars too, but the fully autonomous AI self-driving car will receive the most attention in this discussion.

Here’s the usual steps involved in the AI driving task:

  • Sensor data collection and interpretation
  • Sensor fusion
  • Virtual world model updating
  • AI action planning
  • Car controls command issuance

I’ve mentioned earlier in this discussion that one aspect to be careful about involves the potential of adversarial examples to confuse or mislead the AI system.

Object Transplants

There’s another somewhat similar potential difficulty involving what is sometimes called object transplants.

An interesting research study undertaken by researchers at York University and the University of Toronto provides an insightful analysis of the concerns related to object transplanting (the study was cleverly entitled “The Elephant in the Room”).

Object transplanting can be likened to my earlier comments about the gorilla in the video, though with a slightly different spin involved.

Imagine if you were watching the video that had a gorilla in it.

Suppose that you actually noticed the gorilla when it came into the scene.

If the scene consisted of people tossing a ball back-and-forth, would you be more likely to believe that it was a real gorilla or more likely to believe it is a fake gorilla (i.e., someone in a gorilla suit)?

Assuming that the people kept tossing the ball and did not get freaked out by the presence of the gorilla, I’d bet that you’d mentally quickly deduce that it must be a fake gorilla.

Your context of the video would remain pretty much the same as prior to the introduction of the gorilla.

The appearance of the gorilla did not substantially alter what you thought the scene consisted of.

Would the introduction of the gorilla cause you to suddenly believe that the people must be in the jungle someplace?

Probably not.

Would the gorilla cause you to start looking at other objects in the room and begin to think those might be gorilla related objects?

For example, suppose there was in the room a yellow colored stick.

Before the gorilla appeared, you noticed the stick and just assumed it was nothing more than a stick. Once the gorilla arrived, if you are now shifting mentally and thinking about gorillas, maybe the yellow stick now seems like it might be a banana. You know that gorillas like bananas. Therefore, something that has a somewhat appearance of a banana, might indeed be a banana.

I realize you might scoff at the idea that you would suddenly interpret a yellow stick to be a banana simply because of the gorilla being there. A child watching the video might be more susceptible to making that kind of mental leap. The child has perhaps not seen as many bananas in their lifetime as you have, and thus a yellow stick might seem visually close enough to the resemblance of a banana that a child would mistake it for such. The child didn’t think it was a banana before seeing the gorilla, it was the gorilla that caused the child to re-interpret the scene and the stick-like object in the context of a gorilla being there.

Visual object transplanting can impact the detection aspects of a trained machine learning system such as a convolutional deep neural network in a potentially similar way.

Using the popular Tensorflow object detection capability, and when combined with the Microsoft MS-COCO dataset, the researchers used a picture of a human sitting in a room and playing video games, and then did some object transplanting into the picture to see what the neural network would report about the objects in the picture (essentially, the researchers were doing a Photoshop-style transformation to the picture).

They transplanted an image of an elephant so that it appears in the picture with the sitting human.

In some instances, the neural network did not detect that the elephant was in the picture, presumably not even noticing that it was there (thus, the clever titling of the research study as dealing with the elephant in the room!).

Depending upon where the elephant was positioned in the picture, the neural network at one point reported that the elephant was actually a chair.

In another instance, the elephant was placed near other objects that had been earlier identified, such as cup and a book, and yet the neural network no longer reported having found the cup or the book. There were also instances of switched identifies, wherein the neural network had identified a chair and a couch, but with the elephant nearby to those areas of the picture, the neural network then reported that the chair was a couch and the couch was a chair.

You might complain about this experiment and say that it is perhaps “unfair” to suddenly place an elephant into a picture that has nothing to do with elephants.

The neural network had not been explicitly trained to have a co-occurrence of an elephant and the sitting human playing a video game. Well, the researchers considered this aspect and repeated the experiment but used a picture for which they merely took items already in the picture and moved those selected items around the scene. Once again, there were various inappropriate results produced by the neural network involving object misidentifications of one kind or another.

I would also suggest that we should decidedly not have much sympathy for the neural network per se and the aspect that it had not been trained on the co-occurrence possibilities – it’s inherent inability to readily cope with the co-occurrence aspects “on the fly” so to speak is a weakness that we must overcome overall for such AI systems.

AI Self-Driving Cars And Object Transplants

On a related note, I’ve previously mentioned in the realm of AI self-driving cars that there has been an ongoing debate related to the same notion of object transplanting, specifically the topic of a man on a pogo stick that suddenly appears in the street and near to an AI self-driving car.

For my article that discusses the pogo stick matter, see: https://aitrends.com/selfdrivingcars/egocentric-design-and-ai-self-driving-cars/

There are some AI developers that have argued that it’s understandable that the AI of a self-driving car might not recognize a man on a pogo stick that’s in the street.

By recognizing, I mean that the visual images captured by the self-driving car are examined by the AI and that the AI system was not able to discern that the object in the street consisted of a man on a pogo stick. It detected that an object was there, and had a rather irregular shape, but it was not able to discern that the shape consisted of a person and a pogo stick (in this instance, the two are combined, since the man was on a pogo stick and pogoing).

Why would it be useful or important to discern that the shape consists of a person on a pogo stick?

You, as a thinking human being, and assuming that you’ve seen a pogo stick before, and one that’s in use, you likely know that it involves going up-and-down and also moving forward or backward or side-to-side. If you were driving along and suddenly saw a person pogoing in the street, you’d likely be cognizant that you should be watching out for the potential erratic moves of the pogo stick and its occupant. You could potentially even predict which way the person was going to go, by watching their angle and how hard they were pogoing.

An AI system that merely construes the pogoing human as a blob would not readily be able to predict the behavior of the blob. Predictions are crucial when you drive a car. Human drivers are continually looking around the surroundings of the car and trying to predict what might happen next. That bike rider in the bike lane might be weaving just enough that you are able to predict that they will swerve into your lane, and so you take precautionary measures in-advance. We must expect AI self-driving cars, and especially the true Level 5 self-driving cars, must be able to do this same kind of predictive modeling.

The range of potential problems associated with object transplanting woes includes:

  • In the case of object transplanting, there is the chance that the transplanted object is not detected at all, even though it might normally have been detected in some other context.
  • Or, the confidence level or probability attached to the object certainty might be lessened in comparison to what it might otherwise have been (in the case of the elephant added into the picture and the subsequent missing cup or book, it could be that the neural network had detected the cup and the book but had assigned a very low probability to their identities, and so reported that they weren’t there, based on some threshold level required to be considered present in the picture).
  • The detection of the transplanted object, if detection does occur, might lead to misidentification of other objects in the scene.
  • Other objects might no longer be detected.
  • Or, those other objects might have a lessened probability assigned to them as identifiable objects. There can be both local and non-local effects due to the transplanted object.
  • Other objects might get switched in terms of their identities, due to the introduction of the transplanted object.

Conclusion

For AI self-driving cars, there are a myriad of sensors that collect data about the world surrounding the self-driving car. This includes cameras that capture pictures and video, it includes radar, it includes sonic, it includes LIDAR, and so on. The AI needs to examine the data and try to ferret out what the data indicates about the surrounding objects.

Are those cars ahead of the self-driving car or are they motorcycles?

Are there pedestrians standing at the curb or just a fire hydrant and a light post?

These are crucial determinations for the AI self-driving car and its ability to perform the driving task.

AI developers need to take into account the limitations and considerations that arise due to object transplanting. The AI systems of the self-driving car need to be shaped in a manner that they can sufficiently and safely deal with object transplantation and do so in real-time while the self-driving car is in motion.  The scenery around the self-driving car will not always be pristine and devoid of unusual or seemingly out-of-context objects.

When I was a professor, each year a circus came to town and the circus animals arrived via train, which happened to get parked near the campus for the time period that the circus was in town. A big parade even occurred involving the circus performers marching the animals from next to the campus and over to the nearby convention center. It was quite an annual spectacle to observe.

I mention this because among the animals were elephants, along with giraffes and other “wild” animals. Believe it or not, on the morning of the annual parade, I would usually end-up driving my car right near to the various animals as I was navigating my way onto campus to teach classes for the day. It was as though I had been transported to another world.

If I was using an AI self-driving car, one wonders what the AI might have construed of the elephants and giraffes that were next to the car. Would the AI have suddenly changed context and assumed I was now driving in the jungle? Would it get confused and believe that the light poles were actually tall jungle trees?

I say this last aspect about the circus in some jest but do want to be serious about the facet that it is important to realize the existing limitations of various machine learning algorithms and artificial neural network techniques and tools. AI self-driving car makers need to be on their toes to prepare for and contend with object transplants.

And that’s no elephant joke.

That’s the elephant in the room and on the road ahead for AI self-driving cars.

Copyright 2020 Dr. Lance Eliot

This content is originally posted on AI Trends.

[Ed. Note: For reader’s interested in Dr. Eliot’s ongoing business analyses about the advent of self-driving cars, see his online Forbes column: https://forbes.com/sites/lanceeliot/]

How Businesses Are Making Their Sites More Accessible for 2020

If you take a look around, you’ll notice that parts of our world operate in a way to promote inclusivity for people with disabilities. Buildings often have ramp entrances, guide rails, and elevators for visitors that have limited mobility. Signs and plaques outside of bathrooms have braille writing for the visually impaired.

The reason why many businesses have these types of features is due to the Americans with Disabilities Act - which was passed to provide accessibility for everyone. However, no such law has been passed for online businesses – yet.

The current legislation put into place is known as section 508 compliance. This requires that any type of website or online technology published by an American federal agency must meet specific online accessibility standards. Other countries are following suit with similar legislation.

But just because you are not legally required to create an accessible website – it doesn’t mean that you shouldn’t! Accessibility is more important than ever before, especially considering that 71% of disabled users will exit a website if it is not fully accessible to them.

However, this same study also found that 80% of disabled users also stated that they would choose to purchase from a website with the highest ...


Read More on Datafloq

Industry bodies recommend blockchain policies to be based on its functions

The National e-Governance Division (NeGD), under the Ministry of Electronics and Information Technology (MeitY), had in July 2019 tasked NISG with preparing the policy.

India takes its digital success stories global

While some countries are reaching out on their own seeking help with building platforms, in other cases India is leading the efforts,the official said, adding protocols are being developed for the partnership.

The Data of Self-Tracking: What’s the Future of the Quantified Self?

They seem to be everywhere nowadays: sleep trackers and step trackers, mood monitors and heart rhythm recorders. Thanks to the advent of wearable tech and the ever-expanding Internet of Things (IoT), you can now practically encapsulate your entire life story in the numbers taken from your mobile devices. There’s even a term for this nearly incessant self-monitoring: it’s called “the quantified self,” and it’s changing the way we understand ourselves, our lives, and the way we want to live. But what does all of this relentless pursuit of self-knowledge really mean? What are the benefits and the harms of distilling your life down into a collection of data?What is the “Quantified Self”?Basically, the quantified self (QS) refers to a strategy for tracking a person’s vital data across time in order to identify important patterns. This can include everything from monitoring your biorhythms to tracking your daily activities, whether at home or at work. Most people who engage in QS claim that they want to use their data for specific, actionable purposes. Their goal, ultimately, is to identify opportunities to change their behavior in order to live healthier, happier, and more productive lives.The Role of Wearable Tech in QSQS has actually ...


Read More on Datafloq

Monday, 27 January 2020

Fine-Grained Time Series Forecasting At Scale With Facebook Prophet And Apache Spark

Try this time series forecasting notebook in Databricks

Advances in time series forecasting are enabling retailers to generate more reliable demand forecasts. The challenge now is to produce these forecasts in a timely manner and at a level of granularity that allows the business to make precise adjustments to product inventories. Leveraging Apache Spark™ and Facebook Prophet, more and more enterprises facing these challenges are finding they can overcome the scalability and accuracy limits of past solutions.

In this post, we’ll discuss the importance of time series forecasting, visualize some sample time series data, then build a simple model to show the use of Facebook Prophet. Once you’re comfortable building a single model, we’ll combine Prophet with the magic of Apache Spark™ to show you how to train hundreds of models at once, allowing us to create precise forecasts for each individual product-store combination at a level of granularity rarely achieved until now.

Accurate and timely forecasting is now more important than ever

Improving the speed and accuracy of time series analyses in order to better forecast demand for products and services is critical to retailers’ success. If too much product is placed in a store, shelf and storeroom space can be strained, products can expire, and retailers may find their financial resources are tied up in inventory, leaving them unable to take advantage of new opportunities generated by manufacturers or shifts in consumer patterns. If too little product is placed in a store, customers may not be able to purchase the products they need. Not only do these forecast errors result in an immediate loss of revenue to the retailer, but over time consumer frustration may drive customers towards competitors.

New expectations require more precise time series models and forecasting methods

For some time, enterprise resource planning (ERP) systems and third-party solutions have provided retailers with demand forecasting capabilities based upon simple time series models. But with advances in technology and increased pressure in the sector, many retailers are looking to move beyond the linear models and more traditional algorithms historically available to them.

New capabilities, such as those provided by Facebook Prophet, are emerging from the data science community, and companies are seeking the flexibility to apply these machine learning models to their time series forecasting needs.

Facebook Prophet logo

This movement away from traditional forecasting solutions requires retailers and the like to develop in-house expertise not only in the complexities of demand forecasting but also in the efficient distribution of the work required to generate hundreds of thousands or even millions of machine learning models in a timely manner. Luckily, we can use Spark to distribute the training of these models, making it possible to predict not just overall demand for products and services, but the unique demand for each product in each location.

Visualizing demand seasonality in time series data

To demonstrate the use of Prophet to generate fine-grained demand forecasts for individual stores and products, we will use a publicly available data set from Kaggle. It consists of 5 years of daily sales data for 50 individual items across 10 different stores.

To get started, let’s look at the overall yearly sales trend for all products and stores. As you can see, total product sales are increasing year over year with no clear sign of convergence around a plateau.

Sample Kaggle retail data used to demonstrate the combined fine-grained demand forecasting capabilities of Prophet and SparkNext, by viewing the same data on a monthly basis, we can see that the year-over-year upward trend doesn’t progress steadily each month. Instead, we see a clear seasonal pattern of peaks in the summer months, and troughs in the winter months. Using the built-in data visualization feature of Databricks Collaborative Notebooks, we can see the value of our data during each month by mousing over the chart.

At the weekday level, sales peak on Sundays (weekday 0), followed by a hard drop on Mondays (weekday 1), then steadily recover throughout the rest of the week.

Demonstrating the difficulty of accounting for seasonal patterns with traditional time series forecasting methods

Getting started with a simple time series forecasting model on Facebook Prophet

As illustrated in the charts above, our data shows a clear year-over-year upward trend in sales, along with both annual and weekly seasonal patterns. It’s these overlapping patterns in the data that Prophet is designed to address.

Facebook Prophet follows the scikit-learn API, so it should be easy to pick up for anyone with experience with sklearn. We need to pass in a 2 column pandas DataFrame as input: the first column is the date, and the second is the value to predict (in our case, sales). Once our data is in the proper format, building a model is easy:

import pandas as pd
from fbprophet import Prophet

# instantiate the model and set parameters
model = Prophet(
    interval_width=0.95,
    growth='linear',
    daily_seasonality=False,
    weekly_seasonality=True,
    yearly_seasonality=True,
    seasonality_mode='multiplicative'
)

# fit the model to historical data
model.fit(history_pd)

Now that we have fit our model to the data, let’s use it to build a 90 day forecast. In the code below, we define a dataset that includes both historical dates and 90 days beyond, using prophet’s make_future_dataframe method:

future_pd = model.make_future_dataframe(
    periods=90,
    freq='d',
    include_history=True
)

# predict over the dataset
forecast_pd = model.predict(future_pd)

That’s it! We can now visualize how our actual and predicted data line up as well as a forecast for the future using Prophet’s built-in .plot method. As you can see, the weekly and seasonal demand patterns we illustrated earlier are in fact reflected in the forecasted results.

predict_fig = model.plot(forecast_pd, xlabel='date', ylabel='sales')
display(fig)

This visualization is a bit busy. Bartosz Mikulski provides an excellent breakdown of it that is well worth checking out. In a nutshell, the black dots represent our actuals with the darker blue line representing our predictions and the lighter blue band representing our (95%) uncertainty interval.

Training hundreds of time series forecasting models in parallel with Prophet and Spark

Now that we’ve demonstrated how to build a single model, we can use the power of Apache Spark to multiply our efforts. Our goal is to generate not one forecast for the entire dataset, but hundreds of models and forecasts for each product-store combination, something that would be incredibly time consuming to perform as a sequential operation.

Building models in this way could allow a grocery store chain, for example, to create a precise forecast for the amount of milk they should order for their Sandusky store that differs from the amount needed in their Cleveland store, based upon the differing demand at those locations.

How to use Spark DataFrames to distribute the processing of time series data

Data scientists frequently tackle the challenge of training large numbers of models using a distributed data processing engine such as Apache Spark. By leveraging a Spark cluster, individual worker nodes in the cluster can train a subset of models in parallel with other worker nodes, greatly reducing the overall time required to train the entire collection of time series models.

Of course, training models on a cluster of worker nodes (computers) requires more cloud infrastructure, and this comes at a price. But with the easy availability of on-demand cloud resources, companies can quickly provision the resources they need, train their models, and release those resources just as quickly, allowing them to achieve massive scalability without long-term commitments to physical assets.

The key mechanism for achieving distributed data processing in Spark is the DataFrame. By loading the data into a Spark DataFrame, the data is distributed across the workers in the cluster. This allows these workers to process subsets of the data in a parallel manner, reducing the overall amount of time required to perform our work.

Of course, each worker needs to have access to the subset of data it requires to do its work. By grouping the data on key values, in this case on combinations of store and item, we bring together all the time series data for those key values onto a specific worker node.

store_item_history
    .groupBy('store', 'item')
    # . . .

We share the groupBy code here to underscore how it enables us to train many models in parallel efficiently, although it will not actually come into play until we set up and apply a UDF to our data in the next section.

Leveraging the power of pandas user-defined functions (UDFs)

With our time series data properly grouped by store and item, we now need to train a single model for each group. To accomplish this, we can use a pandas User-Defined Function (UDF), which allows us to apply a custom function to each group of data in our DataFrame.

This UDF will not only train a model for each group, but also generate a result set representing the predictions from that model. But while the function will train and predict on each group in the DataFrame independent of the others, the results returned from each group will be conveniently collected into single resulting DataFrame. This will allow us to generate store-item level forecasts but present our results to analysts and managers as a single output dataset.

As you can see in the abbreviated code below, building our UDF is relatively straightforward. The UDF is instantiated with the pandas_udf method which identifies the schema of the data it will return and the type of data it expects to receive. Immediately following this, we define the function that will perform the work of the UDF.

Within the function definition, we instantiate our model, configure it and fit it to the data it has received. The model makes a prediction, and that data is returned as the output of the function.

@pandas_udf(result_schema, PandasUDFType.GROUPED_MAP)
def forecast_store_item(history_pd):

    # instantiate the model, configure the parameters
    model = Prophet(
        interval_width=0.95,
        growth='linear',
        daily_seasonality=False,
        weekly_seasonality=True,
        yearly_seasonality=True,
        seasonality_mode='multiplicative'
    )

    # fit the model
    model.fit(history_pd)

    # configure predictions
    future_pd = model.make_future_dataframe(
        periods=90,
        freq='d',
        include_history=True
    )

    # make predictions
    results_pd = model.predict(future_pd)

    # . . .

    # return predictions
    return results_pd

Now, to bring it all together, we use the groupBy command we discussed earlier to ensure our dataset is properly partitioned into groups representing specific store and item combinations. We then simply apply the UDF to our DataFrame, allowing the UDF to fit a model and make predictions on each grouping of data.

The dataset returned by the application of the function to each group is updated to reflect the date on which we generated our predictions. This will help us keep track of data generated during different model runs as we eventually take our functionality into production.

from pyspark.sql.functions import current_date

results = (
    store_item_history
    .groupBy('store', 'item')
    .apply(forecast_store_item)
    .withColumn('training_date', current_date())
    )

Next steps

We have now constructed a forecast for each store-item combination. Using a SQL query, analysts can view the tailored forecasts for each product. In the chart below, we’ve plotted the projected demand for product #1 across 10 stores. As you can see, the demand forecasts vary from store to store, but the general pattern is consistent across all of the stores, as we would expect.

Demand Forecasting Projections Visualized Using Databricks Notebook Visualizations

As new sales data arrives, we can efficiently generate new forecasts and append these to our existing table structures, allowing analysts to update the business’s expectations as conditions evolve.

--

Try Databricks for free. Get started today.

The post Fine-Grained Time Series Forecasting At Scale With Facebook Prophet And Apache Spark appeared first on Databricks.