Productionizing Machine Learning with Delta Lake

Posted Leave a commentPosted in AI, Apache Spark, Company Blog, Data Engineering, Delta Lake, Ecosystem, Education, Engineering Blog, Machine Learning, Platform

Try out this notebook series in Databricks – part 1 (Delta Lake), part 2 (Delta Lake + ML) For many data scientists, the process of building and tuning machine learning models is only a small portion of the work they do every day. The vast majority of their time is spent doing the less-than-glamorous (but […]

Simplifying Streaming Stock Analysis using Delta Lake and Apache Spark: On-Demand Webinar and FAQ Now Available!

Posted Leave a commentPosted in ACID Transactions, Apache Spark, Company Blog, Delta Lake, Education, Engineering Blog, Financial Services, Product, Streaming, Structured Streaming, Time Travel, Unified Batch and Streaming Sync

On June 13th, we hosted a live webinar — Simplifying Streaming Stock Analysis using Delta Lake and Apache Spark — with Junta Nakai, Industry Leader – Financial Services at Databricks, John O’Dwyer, Solution Architect at Databricks, and Denny Lee, Technical Product Marketing Manager at Databricks. This is the first webinar in a series of financial […]

Detecting Bias with SHAP – The Databricks Blog

Posted Leave a commentPosted in Apache Spark, Bias, Deep Learning, Education, Engineering Blog, Machine Learning, MLflow, SHAP, Stack Overflow

StackOverflow’s annual developer survey concluded earlier this year, and they have graciously published the (anonymized) 2019 results for analysis. They’re a rich view into the experience of software developers around the world — what’s their favorite editor? how many years of experience? tabs or spaces? and crucially, salary. Software engineers’ salaries are good, and sometimes […]

New videos from Databricks Academy: Introduction to Natural Language Processing—Latent Semantic Analysis

Posted Leave a commentPosted in Announcements, Company Blog, Education

Databricks’ commitment to education is at the center of the work we do. Through Instructor-Led Training, Certification, and Self-Paced Training, Databricks Academy provides strong pathways for users to learn Apache Spark™ and Databricks to push their knowledge to the next level. Our latest offering is a series of short videos introducing the Natural Language Processing […]

Efficient Databricks Deployment Automation with Terraform

Posted Leave a commentPosted in CI/CD, cloud automation, Company Blog, Customers, Ecosystem, Education, Engineering Blog, Platform

Managing cloud infrastructure and provisioning resources can be a headache that DevOps engineers are all too familiar with. Even the most capable cloud admins can get bogged down with managing a bewildering number of interconnected cloud resources – including data streams, storage, compute power, and analytics tools. Take, for example, the following scenario: a customer […]

Detecting Financial Fraud at Scale with Decision Trees and MLflow on  Databricks

Posted Leave a commentPosted in Apache Spark, Company Blog, Decision tree, Education, Engineering Blog, financial, Financial Markets, Financial Services, Fraud, Fraud Detection, Machine Leanring, Machine Learning, Platform

Try this notebook in Databricks Detecting fraudulent patterns at scale is a challenge, no matter the use case. The massive amounts of data to sift through, the complexity of the constantly evolving techniques, and the very small number of actual examples of fraudulent behavior are comparable to finding a needle in a haystack while not […]

Understanding Dynamic Time Warping – The Databricks Blog

Posted Leave a commentPosted in Apache Spark, Company Blog, Dynamic Time Warping, Education, Engineering Blog, Machine Learning, Platform

Try this notebook in Databricks This blog is part 1 of our two-part series Using Dynamic Time Warping and MLflow to Detect Sales Trends. To go to part 2, go to Using Dynamic Time Warping and MLflow to Detect Sales Trends. The phrase “dynamic time warping,” at first read, might evoke images of Marty McFly […]

Using Dynamic Time Warping and MLflow to Detect Sales Trends

Posted Leave a commentPosted in Apache Spark, Company Blog, Dynamic Time Warping, Education, Engineering Blog, Machine Learning, MLflow, Platform

Try this notebook series in Databricks This blog is part 2 of our two-part series Using Dynamic Time Warping and MLflow to Detect Sales Trends.  The phrase “dynamic time warping,” at first read, might evoke images of Marty McFly driving his DeLorean at 88 MPH in the Back to the Future series. Alas, dynamic time warping does […]

Koalas: Easy Transition from pandas to Apache Spark

Posted Leave a commentPosted in Announcements, Apache Spark, Company Blog, Data Science, Ecosystem, Education, Engineering Blog, Machine Learning, Open Source, Pandas, python

Today at Spark + AI Summit, we announced Koalas, a new open source project that augments PySpark’s DataFrame API to make it compatible with pandas. Python data science has exploded over the past few years and pandas has emerged as the lynchpin of the ecosystem. When data scientists get their hands on a data set, […]

University Web Developer Programs Must Prep Students For Big Data Era

Posted Leave a commentPosted in Big Data, Education, News, SmartData Collective Exclusive, university, web developer programs, web developers

Big data is playing a monumental role in economic development in 2019. It is being used in every industry from healthcare delivery to insurance modeling to political science. According to economic analysts, the market size for big data is currently estimated to be $203 billion. Unfortunately, there is a growing shortage of data scientists. As […]