State of Fast Feedback in Data Science Projects

Let’s talk about the productivity in Data Science and Machine Learning projects (DSML)

DSML projects can be quite different from the software projects: a lot of R&D in a rapidly evolving landscape, working with data, distributions and probabilities instead of code. However, there is one thing in common: iterative development process matters a lot.

For example in software engineering, rapid iterations help a lot in debugging complex issues or working towards a tricky issue. In product development, an ability to rapidly roll out new version can be a deal-breaker for achieving customer satisfaction. Paul Graham eloquently covers that in his “Beating the averages” essay.

Likewise, in Data Science and Machine Learning projects, iterations help data scientists to rapidly test their theories and converge towards the solution that will create value. If we assume that 87% of data science projects fail (which looks about right to me), then having a fast feedback loop could help to get to the successful 13% faster.

Yet, there is a problem in the industry with that.

Let use a basic data science pipeline as an example. It will have a predefined structure to make it easier for collaboration between different teams in the department.

There will be the following steps:

  1. Initialise the pipeline run, deriving any per-run variables from the initial config

  2. Load and prepare the training data

  3. Perform model training

  4. Evaluate the model on a separate dataset

  5. Prepare the model for the use

  6. Run batch prediction against the resulting model

The de facto language for the pipelines in Python. We can provide a minimal implementation in a console application and run it locally. On my laptop it takes ~0.3-0.5 sec.

 

That is good enough.

If the computation overhead of a real pipeline is 5 minutes, then we could run up to 12 iterations in an hour.

However, the industry way to run these pipelines is via Kubeflow (ML toolkit for Kubernetes). Google Vertex is one of the most stable implementations.

If we map our pipeline components to a Kubeflow pipeline, we’ll get something like that:

 

How many experiments per our can we run here?

At this point, the computation overhead doesn’t even matter. Since it takes 33 minutes per run, we could run only up to experiment per hour.

The execution takes 5000x more time on Vertex than it takes on a local machine. Although that time is a paid compute time, the biggest hit is not a financial one, but more of a productivity loss.

And that is the most frustrating problem with the state of the data science pipelines today. Major hosting players make more money from less efficient data science pipelines. This might reduce incentives to prioritize performance-improving changes. This in turn negatively impacts the ability of small data science teams to have fast feedback loops and innovate efficiently.

Navigationsbild zu Data Science
Service

AI & Data Science

We offer comprehensive solutions in the fields of data science, machine learning and AI that are tailored to your specific challenges and goals.

Businessmen work with stock market investments using smartphones to analyze trading data. smartphone with stock exchange graph on screen. Financial stock market
Service

Data Management & Data Science Consulting

With our data consulting solutions, we help you unlock the full potential of your data: from analysis and forecasting to process optimisation and successful AI projects

Articifial Intelligence & Data Science
Service

Artificial Intelligence & Data Science

Data Science is all about extracting valuable information from structured and unstructured data.

Mitarbeiterin arbeitet am Laptop
Asset

novaPredict for Actuarial Data Science

Our solution enables insurers to run complex actuarial models in a scalable, transparent, and cost-effective manner, without relying on rigid off-the-shelf software.

Referenz

Jira Integration of Demand and Project Portfolio Management

In the area of demand and project portfolio management, catworkx was also able to demonstrate the great flexibility of Jira in a customer project and show that relevant business data and influencing..

2023 Referenz IAM Teaserbild SID
Success Story

Elimination of data clutter through identity management.

This is how we create transparency, efficiency, and say goodbye to data clutter. Find out how our experts are making history with SID. Read more about it.

Wissen 4/14/23

General Data Protection Regulation of idea management

Walldorf-based dacuro GmbH provides the external data protection officer for companies, helps with the fulfillment of documentation obligations and advises on all aspects of data protection. Fulfilling the requirements of the GDPR without blocking everyday life is the claim of dacuro GmbH. The team of lawyers and IT specialists provides support for all GDPR challenges, whether they are of a legal or technical nature.

Lösung 9/21/22

Portfolio Project Management (PPM)

How Project Portfolio Management with Atlassian Tools supports global project and QM tasks including Cross-Project Knowledge Management.

Wissen 5/2/24

Unlock the Potential of Data Culture in Your Organization

Are you ready to revolutionize your organization's potential by unleashing the power of data culture? Imagine a workplace where every decision is backed by insights, every strategy informed by data, and every employee equipped to navigate the digital landscape with confidence. This is the transformative impact of cultivating a robust data culture within your enterprise.

Referenz

Flexibility in the data evaluation of a theme park

With the support of TIMETOACT, an theme park in Germany has been using TM1 for many years in different areas of the company to carry out reporting, analysis and planning processes easily and flexibly.

Launch (ESA)
Success Story

ESA: Data Factory, the Single Source of Truth

With the Data Factory, the European Space Agency (ESA) has created a single source of truth – ensuring transparent data and project status, more efficient processes and sustainable decision-making.

Beladenes Containerschiff auf dem Meer
Success Story

dteq: Web-based project management

Off-the-shelf software does not always meet all requirements, whilst bespoke development can cause costs to skyrocket. With the Business Productivity Framework, the dteq group has found the solution.

Mann am Arbeitsplatz mit PC
Service

First Level Support for fast IT support

We provide First Level Support for your business, prioritize tickets efficiently and relieve your IT team in day-to-day operations. Personal, efficient and reliable.

Headerbild Data Insights
Service

Data Insights

With Data Insights, we help you step by step with the appropriate architecture to use new technologies and develop a data-driven corporate culture

Easy Cloud Solution
Produkt

Big Data

Extract valuable information from data - Take advantage of serverless, integrated end-to-end data analytics services to leave traditional limitations behind.

KnowledgeBase

Data Privacy

CLOUDPILOTS Data Privacy

Blog 9/27/22

Creating solutions and projects in VS code

In this post we are going to create a new Solution containing an F# console project and a test project using the dotnet CLI in Visual Studio Code.

Referenz

Integrated Project and User Portal (IPUP)

For an automotive client, catworkx developed a tool for Jira Service Management that enables automated setup of projects and transparent user assignment – flexible and scalable.

Logo Armacell
Referenz

Bundled expertise for fast mail migration to M365

Together with novaCapta, TIMETOACT supports Armacell as a Managed Service Partner for a successful mail migration ► Read Success Story now

Use Case

Use Case: Müller-BBM - faster Expertise with AI

Semantic search across documents & sites enables Müller-BBM to find data in seconds – use knowledge, don’t hunt for it.

Bleiben Sie mit dem TIMETOACT GROUP Newsletter auf dem Laufenden!