Smarter Week

How to automate it

How to automate “prepare features and training datasets”

Here are 2 ways to spend less time on this, best first. Each comes with steps you can follow today and, for AI fixes, a prompt to copy.

120 min
typically, once a week
50%
of the time can be automated
Easy
to set up

Fix 1 of 2

AIBest fix

Clean and explore data in a notebook with AI built in

Hex, Databricks notebooks, Jupyter AI, Colab with Gemini and ChatGPT or Claude data analysis write the cleaning and reshaping code from plain-English instructions. You check the output.

Typically saves about 40% of the time10 min to set up
  1. 1Open the data in an AI-enabled notebook (Hex, Databricks, Colab, Jupyter AI), or upload a file to an approved assistant.
  2. 2Describe the cleaning you need using the prompt below.
  3. 3Read the generated code; check row counts before and after each step.
  4. 4Save it as a reusable notebook or move it upstream into the pipeline.
Prompt to copy
I have a dataset with these columns: [LIST COLUMNS AND TYPES]. Write [PANDAS / POLARS / SQL] code to: [e.g. standardize dates to ISO, trim and lowercase emails, remove exact duplicates, split full name into first and last, flag rows with missing customer IDs]. Print row counts before and after each step, and list any assumptions. Do not drop rows silently.

Tools: Hex Magic · Databricks Assistant · Jupyter AI · Google Colab with Gemini · ChatGPT or Claude data analysis

Fix 2 of 2

AI

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

Quick wins

Have you tried…

Have you used an AI-enabled notebook to write data cleaning or analysis code from plain English?
Hex, Databricks, Colab and Jupyter AI write pandas or SQL code from a description of what you want, and you check the results step by step.

Who does this task

Roles in our library that list this as one of their common tasks. Each guide covers the rest of that role’s week.

HourLeak · the 8-minute work audit

How many hours does this cost you?

The free 8-minute check works out where your week goes and gives you your top fixes. The team scan does the same for everyone and adds it up, so you know which leaks to fix first.

Answers are anonymous. Leaders only see team totals.

Other common tasks for Data scientists