| Dataset requirements specification | Purpose, users, scope, records, labels, quality, constraints and acceptance criteria | Document and requirements register | Discovery | Use-case and risk decisions | Product or model owner |
| Source and eligibility assessment | Candidate sources, rights, accessibility, limitations, freshness and suitability | Assessment matrix | Assessment | Source access and policies | Data owner |
| Schema and ontology | Fields, types, entities, relationships, labels and metadata definitions | Data dictionary, diagrams or machine-readable schema | Design | Subject-matter validation | Data architect or lead scientist |
| Sampling and split plan | Population, strata, proportions, edge cases, exclusions and independence rules | Plan and coverage matrix | Design | Population evidence | Data science lead |
| Annotation specification | Label definitions, instructions, examples, ambiguity handling and QA | Guideline and decision tree | Design | Expert adjudication | Annotation or domain lead |
| Quality and control framework | Validation rules, thresholds, provenance, versioning, access and monitoring | Control matrix and test catalogue | Validation | Risk and policy review | Data governance owner |
| Implementation backlog | Build tasks, dependencies, priorities, owners, decisions and acceptance tests | Roadmap or work-item register | Handover | Delivery capacity and priorities | Programme or engineering lead |