Task 1 · typically 60 min, a few times a week
Clean and reshape messy data
Clean and explore data in a notebook with AI built in
Hex, Databricks notebooks, Jupyter AI, Colab with Gemini and ChatGPT or Claude data analysis write the cleaning and reshaping code from plain-English instructions. You check the output.
- 1Open the data in an AI-enabled notebook (Hex, Databricks, Colab, Jupyter AI), or upload a file to an approved assistant.
- 2Describe the cleaning you need using the prompt below.
- 3Read the generated code; check row counts before and after each step.
- 4Save it as a reusable notebook or move it upstream into the pipeline.
I have a dataset with these columns: [LIST COLUMNS AND TYPES]. Write [PANDAS / POLARS / SQL] code to: [e.g. standardize dates to ISO, trim and lowercase emails, remove exact duplicates, split full name into first and last, flag rows with missing customer IDs]. Print row counts before and after each step, and list any assumptions. Do not drop rows silently.
Tools: Hex Magic · Databricks Assistant · Jupyter AI · Google Colab with Gemini · ChatGPT or Claude data analysis
One more way to fix itHide the other fixes
Fix the data at the source instead of cleaning it every time
If you clean the same mess every week, the fix belongs upstream: a required field, a dropdown instead of free text, or a cleaning step in the pipeline.
- 1List the cleaning steps you repeat most.
- 2For each, find where the bad data enters (a form, a CRM field, an import).
- 3Ask the system owner for validation, dropdowns or required fields.
- 4Move any cleaning that must stay into a shared pipeline model so nobody does it by hand.
Tools: Your CRM, ERP or form tool settings · dbt or your pipeline tool