
Begin
14 pages · ~28 min
Documenting Data Science Decisions
This training teaches data scientists how to document decisions effectively, helping teams improve reproducibility, transparency, and collaboration across data science projects.
A digital instructor presents all 14 pages. Hold “Ask” at any point and ask out loud — the answer comes from this course. No sign-up needed.
What you’ll learn
- 01Documenting Decisions in Data ScienceWelcome. This course is about documenting decisions in data science, and treating that documentation as part of the craft, not paperwork we squeeze in at the end. Good records give us reproducibility, auditability, and knowledge retention, and they cut down on rework. Think about a model someone rebuilt from scratch because no one remembered why a key feature was dropped six months earlier. That is the cost we are trying to avoid. Decisions happen across the whole lifecycle: framing the problem, selecting data, choosing methodology, validating results, and deploying. So we document in layers. Quick notes for working context, decision records that capture the choice, the alternatives, and the reasoning, and model, data, or system cards for the durable summary. The course follows a simple path. We start with principles, then templates, then team workflows, and finish with hands-on practice. Keep one idea in mind throughout. The goal is not more documents. It is decisions that remain understandable to someone who was not in the room. Next, let us look at what undocumented decisions actually cost us.
nvlpubs.nist.govairc.nist.govairc.nist.gov+22 min - 02The Cost of Undocumented DecisionsSo let us turn to what undocumented decisions actually cost us. Undocumented assumptions and silent cleaning choices hide the real reasoning behind our results. When those choices go unrecorded, reproductions fail, compliance reviews find gaps, revenue slips, and onboarding slows to a crawl. Documentation debt compounds. Stale or missing records erode trust in every document, even the accurate ones. Consider a real case: training and serving features diverged, and that single gap cost one team one point five million dollars in lost revenue. Now think about AI-generated documentation. Without human insight, it creates volume without trust. Developers spot a few invented references and stop reading altogether. So the takeaway is simple. If a choice shaped your results, write down the why, and keep it current. That way documentation stays a professional asset, not a liability. Next, let us look at the background and regulatory context.
journals-sol.sbc.org.brcs.montana.eduaaltodoc.aalto.fi+21 min - 03Background and Regulatory ContextLet's step back and look at why structured decision records are becoming standard practice. For years, many teams relied on ad hoc notes, comments in notebooks, or memory. That worked at small scale. It does not scale to regulated, high-stakes systems.
The clearest driver is the EU AI Act. Article eleven requires technical documentation for high-risk systems, and Annex IV spells out what that means: design rationale and assumptions, data provenance and labeling, validation and testing procedures, performance metrics, and known limitations. Those obligations apply from August 2026. High-risk rules phase in through 2027, and rules for AI embedded in medical devices and other physical products follow into 2028. Financial institutions already know this pattern from SR 11-7 and SS1/23, which expect documentation detailed enough for an independent third party to understand and replicate a model. The FDA sets similar expectations for health applications. And voluntary frameworks, like the NIST AI Risk Management Framework with its Govern, Map, Measure, and Manage functions, and ISO/IEC 42001, align documentation with responsible AI practice.
So the direction is consistent: write it down, keep it current, and make the reasoning reviewable. Next, we will build the vocabulary that these documents rely on, in Core Concepts and Vocabulary.
ai-act-service-desk.ec.europa.euai-act-service-desk.ec.europa.euai-act-service-desk.ec.europa.eu+22 min - 04Core Concepts and VocabularyLet's ground this in the vocabulary we'll use from here on. First, artifacts. A decision record captures one decision. A decision log collects them. An assumption register tracks what we're taking for granted. Model cards and data cards summarize the model and its data. Second, metadata. Every useful record answers the same questions: what was the context? What options did we weigh? Why did we choose this one? What evidence supported it? Who owns it? When was it written? What's its status? And what condition would make us revisit it? Third, note the distinction between atomic decisions, which stand alone, and composite decisions, which bundle several. And between reversible choices, cheap to undo, and irreversible ones like a schema migration in production, which deserve more scrutiny. Traceability ties it together by linking each decision to data versions, commits, experiments, and model artifacts. Watch for three anti-patterns: post hoc rationalization, writing the story after the fact; decision theater, documentation created for show rather than use; and documentation debt, the accumulated cost of decisions nobody wrote down. When and What to Document.
distilledpatterns.orgcrunchingthedata.comar5iv.labs.arxiv.org+22 min - 05When and What to DocumentNow let's talk about when and what to document. Not everything deserves a record. The skill is knowing which decisions do.
First, watch for decision triggers. Data inclusion, feature engineering, model selection, thresholds, and deployment gating. These are the moments where a choice shapes everything downstream.
When one of those triggers fires, score it on four dimensions. Impact. Reversibility. Uncertainty. And cross-team dependency. High impact, hard to reverse, uncertain, or touching other teams? That's worth writing down.
Document selectively. Record baselines and direction-setting choices. Skip routine steps that anyone could redo in an afternoon. And capture negative results and abandoned paths. That is often the most valuable note you will write, because it stops a colleague from burning a week on the same dead end.
Finally, assign clear roles for each record. Who writes it. Who reviews it. Who approves it. That turns documentation from a personal habit into a team agreement. The test is simple. If someone new joined tomorrow, would this record save them real time?
Next, we will look at the anatomy of a decision record.
distilledpatterns.orgcrunchingthedata.comar5iv.labs.arxiv.org+22 min - 06Anatomy of a Decision RecordLet's look at the anatomy of a decision record. Most records share a standard set of fields: title, status, context, decision, consequences, alternatives, and a review trigger. Keep it short, ideally one page, and link to supporting material rather than inlining it.
Of these, alternatives are the highest-value section. Writing down what you rejected, and why, stops the team from re-litigating the same options six months later. If the rejection reason no longer holds, the record tells you exactly what to revisit.
When you need something compact, use a Y-statement. In context X, facing Y, I decided Z to achieve W, accepting V. One sentence, fully specific, no placeholders.
Record uncertainty honestly. Note your confidence level, open questions, assumptions, and the conditions that should trigger a review. And link your evidence: notebooks, dashboards, tickets, experiment IDs, data versions, and commits. That is what makes a decision reconstructable later.
Next, we turn to Model Cards, Data Cards, and System Cards.
1 min - 07Model Cards, Data Cards, and System CardsNow let's look at the documentation artifacts that describe what a system is. Model cards cover intended use, training data, evaluation metrics, known limitations, and safe-use guidance. Data cards, sometimes called datasheets, go deeper into provenance: how data was collected, its consent basis, and known biases. System cards are the public-facing summary of capabilities, limitations, and key design choices. A useful way to think about it: these artifacts summarize what the system is. Decision records explain why it was built that way. You need both, and they complement each other. One practical note: match your artifact selection to team size, maturity, and regulatory exposure. A small research team may only need a lightweight model card. A production system in a regulated domain likely needs the full set. Start with the artifact that answers the question your reviewers are actually asking. Next, we'll turn to capturing decisions during analysis and experimentation.
ai-act-service-desk.ec.europa.euai-act-service-desk.ec.europa.euai-act-service-desk.ec.europa.eu+22 min - 08Capturing Decisions During Analysis and ExperimentationNow let's look at capturing decisions as they happen. During exploratory analysis, write down your choices, not just your results: which rows or variables you included or excluded, how you cleaned and transformed them, and how you handled missing values through imputation. For each hypothesis, log what you tested, what came back, what you learned, and why the final approach won. When an experiment fails, that is still a decision worth recording. Then link those choices to your experiment trackers, like MLflow or W and B, and to your data catalog, so metrics and reasoning live together. Here's a practical check, borrowed from the analyst and inspector idea: could a reviewer reproduce your workflow from the documentation alone? If not, the workflow is missing important decisions. And if a large language model sits on the data path, record more: the model version, the prompts, the sampling parameters like temperature and seed, and how deterministic the output really is. Next, we'll move into practical strategies for teams.
1 min - 09Practical Strategies for TeamsLet's move from principles to practice, with five strategies you can adopt as a team. First, embed decision records into the rituals you already have. Sprint planning, design reviews, and model sign-off are natural checkpoints. So when a model is promoted to production, the rationale should be captured right there, not reconstructed later.
Second, create one shared decision log, with a consistent location and naming conventions. Settled choices stay queryable, so questions like why didn't we use approach B get answered in seconds rather than in a careful archaeology of old pull requests.
Third, automate the capture. Let commits, run IDs, dataset versions, and registry stage transitions write the evidence for you. A promotion gate becomes an audit record. That reduces effort and makes the trail far harder to lose.
Fourth, scale deliberately. Start with a few champions, then publish playbooks and conventions so the practice spreads as a team standard, not a mandate from above.
Fifth, for legacy projects, write retroactive records. Prioritize by risk. You will not document everything, and you should not try.
Next, let's look at tools and workflow integration.
journals-sol.sbc.org.brcs.montana.eduaaltodoc.aalto.fi+22 min - 10Tools and Workflow IntegrationLet's talk about how to fit decision records into the tools you already use. In code repositories, Markdown decision records living in a docs or decisions folder work well, because they sit next to the code and version alongside it. For non-code decisions that involve stakeholders, a wiki template in Confluence or Notion gives you a shared, living page. Model registries add enforcement: annotations and gated stage transitions mean a model can't move forward without documented rationale. Automated evidence matters too. Capture run IDs, dataset IDs, model hashes, and approval events, so the record links back to what actually happened. Then choose tools by team size, maturity, and regulatory needs. The goal is traceability, not tool sprawl. Next, we'll look at review triggers and revisiting decisions.
1 min - 11Review Triggers and Revisiting DecisionsLet's talk about review triggers and revisiting decisions. Decisions decay. They don't stay valid forever. So when you record a decision, name the trigger that should force a revisit. Good triggers are concrete. A cost threshold. A vendor change. A shift in data availability. For example, revisit your embedding model choice if inference cost passes two thousand dollars a month. And when you do revisit, supersede rather than edit. Link the old record to its replacement. Keep the original intact, because it shows what you believed, and why, at the time. Then watch for context drift. Recheck the original assumptions against current reality. Was that latency budget still two hundred milliseconds? Is the vendor still offering that pricing tier? Finally, record outcomes. Worked, mixed, or reversed. That builds a searchable history of which approaches held up, and where your team tends to reverse itself. Next, we move into quality, audit, and governance.
1 min - 12Quality, Audit, and GovernanceQuality, audit, and governance take us from writing decisions down to making them defensible. Start with quality criteria. Documentation should be complete, clear, traceable, timely, and honest about trade-offs. If you chose a simpler model for stability, say so, and say what you gave up.
Next, peer review and approval. Independent review matters, but so does effective challenge. A reviewer who only checks formatting is not adding much. You want clear sign-off: who approved it, what they reviewed, and what conditions they attached.
Audit readiness raises the bar. A qualified third party with relevant expertise should be able to understand your decisions and replicate your parameter estimation. That is the standard regulators describe, and it is worth designing toward even when no audit is imminent.
Governance models vary. Centralized ownership gives consistency and control. Federated ownership keeps documentation close to the teams and moves faster, but it needs shared templates and periodic checks for drift.
Finally, track effectiveness. Watch onboarding time, how often old decisions get re-litigated, and themes in audit findings. Those measures tell you whether documentation is working, not just present.
That leads us to common pitfalls and how to avoid them.
nvlpubs.nist.govairc.nist.govairc.nist.gov+22 min - 13Common Pitfalls and How to Avoid ThemLet's talk about the pitfalls that quietly undermine decision documentation, and how to counter them.
The first is documentation debt. Missing, outdated, or inconsistent records compound silently. One reviewer estimates a feature based on docs that no longer match reality, and the error propagates. You only notice when it's expensive.
The second is post hoc rationalization. Records written after the fact tend to omit the real trade-offs. You remember the decision, not the alternatives you rejected. Write the record when you decide, not weeks later.
Then there's the opposite failure: over-documentation and documentation theater. Endless templates that slow delivery without adding value. If AI can generate it from the code alone, it probably doesn't belong in a decision record.
Inconsistent formats make search, review, and automation difficult. Six templates across four teams means nobody can compare decisions or find them again.
And culture matters most. Low incentives, time pressure, and fear of scrutiny all push documentation aside. Counter these with lightweight templates, visible leadership support, and recognition for good records. When leads treat the decision log as real work, teams follow. So the takeaway is simple: keep records lightweight, consistent, and written at the moment of choice.
Next, we'll move into the hands-on workshop and your next steps.
journals-sol.sbc.org.brcs.montana.eduaaltodoc.aalto.fi+22 min - 14Hands-On Workshop and Next StepsLet's bring this together with a little practice. Your guided exercise: pick one analytical choice you made recently, maybe a threshold, a feature you dropped, a sampling decision, and write a decision record for it. Keep it short: the question, the options, the reasoning, the owner, and a review date. Then swap records with a colleague. Peer review works best with a quality checklist, and one analyst-inspector reproduction test: can the other person follow your evidence to your decision and carry out the next action without asking you anything? That single question catches most gaps. From there, adapt the templates to your own tools and your regulatory environment, because a regulator will want specific things written down. And build a thirty, sixty, ninety day adoption plan, starting small: one record for your next significant call. Find templates, further reading, and internal champions to keep the habit alive. Thank you for working through this course. Documenting decisions is a craft, and every record you write makes your team's reasoning a little more durable. Go write your first one.
2 min
Take the deck with you
Download this course as a file — free, no sign-up needed.
- PDF handoutEvery slide page, ready to print or share.15 pages · 3.3 MBDownload
- Narrated PowerPointThe deck that presents itself — every slide carries the digital human's narration video.15 pages · 14.2 MBDownload
- PowerPoint slidesThe full deck as a .pptx — open it in PowerPoint, Keynote, or Google Slides.15 pages · 3.2 MBDownload
Free to use in your own training — please keep the PersonWise credit page at the end.
Have your own deck? Turn it into a course
Sources consulted
Web sources consulted while building this course.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0) — nvlpubs.nist.gov
- AI RMF PLAYBOOK — airc.nist.gov
- AI RMF - AIRC — airc.nist.gov
- AI RMF Core - AIRC — airc.nist.gov
- AI Risk Management Framework | NIST — nist.gov
- A Method to Support Documentation Technical Debt Management | iSys - Journal of Information Systems — journals-sol.sbc.org.br
- Hearing the Voice of Software Practitioners on Causes, Effects, and Practices to Deal with Documentation Debt — cs.montana.edu
- Torkki, Säde; Penttinen, Esko; Rinta-Kahila, Tapani; Ruissalo, Joona Recommendations for Dealing with Unexpected Challenges in Paying Back Documentation Debt — aaltodoc.aalto.fi
- Hidden Technical Debt in Machine Learning Systems — papers.neurips.cc
- Documentation Debt Report — doc.holiday
- Annex IV | AI Act Service Desk — ai-act-service-desk.ec.europa.eu
- Article 11: Technical documentation | AI Act Service Desk — ai-act-service-desk.ec.europa.eu
- ANNEX XI | AI Act Service Desk — ai-act-service-desk.ec.europa.eu
- Article 53: Obligations for providers of general-purpose AI models | AI Act Service Desk — ai-act-service-desk.ec.europa.eu
- Implementation Guidance for the EU AI Act — futurium.ec.europa.eu
- Decision Traceability | DistilledPatterns — distilledpatterns.org
- Data science design documents - Crunching the Data — crunchingthedata.com
- [2105.00687] Learning by Design: Structuring and Documenting the Human Choices in Machine Learning Development — ar5iv.labs.arxiv.org
- Practical Interpretability - Page 4 of 5 | OneNoughtOne — onenoughtone.com
- best-practices/MLOps/3a_dataScience_best_practices.ipynb — github.com