Celebrating 25 years of DDD's Excellence and Social Impact.

Author name: kevin sahotsky

Kevin Sahotsky leads strategic partnerships and go-to-market strategy at Digital Divide Data, with deep experience in AI data services and annotation for physical AI, autonomy programs, and Generative AI use cases. He works with enterprise teams navigating the operational complexity of production AI, helping them connect the right data strategy to real model performance. At DDD, Kevin focuses on bridging what organizations need from their AI data operations with the delivery capability, domain expertise, and quality infrastructure to make it happen.

Avatar of kevin sahotsky
AI Governance Frameworks

AI Governance Frameworks: What Boards and C-Suites Need to Own About Data Decisions

Kevin Sahotsky

Here’s a pattern I’ve started seeing in boardrooms: the board asks management whether the company has an AI policy, management says yes, everyone moves to the next agenda item, and the actual decisions that create AI liability keep getting made three levels down, by default, by whoever happens to be assembling training data that week. The policy exists. The governance doesn’t.

A quick word on my vantage point: I lead go-to-market and strategic partnerships at Digital Divide Data, and the change I’ve noticed across this market over the past two years is who shows up to our conversations. It used to be data science leaders. 

Increasingly, the people in those conversations carry legal, risk, and audit responsibility, sometimes one person wearing all three hats, and the questions they bring reflect board priorities trickling down into the programs we work on. The numbers explain why that pressure is only now reaching the working level: in Deloitte’s Global Boardroom Program survey, 45 percent of directors and executives said AI wasn’t on the board agenda at all, and 79 percent said their boards had limited, minimal, or no knowledge or experience with AI. The 2025 follow-up showed the agenda gap narrowing to 31 percent, which means the priorities are starting to cascade, but they’re cascading from boards that mostly can’t yet interrogate the topic.

Here’s the thesis of this piece: when boards do engage with AI, they tend to govern the models and the use cases, because that’s where the demos are. But the least governed liability lives in the data decisions. The board oversight trackers make the point almost by accident: EY’s review of Fortune 100 disclosures and NACD’s annual board survey measure AI committee assignments, agenda time, and risk factor disclosure in detail, and neither contains a category for training data provenance or the data supply chain at all. Gartner’s analysis of why GenAI projects get abandoned after proof of concept lists poor data quality first among the causes, ahead of risk controls and cost, and the EU AI Act writes data governance obligations directly into law for high-risk systems. 

Failures against those obligations carry fines of up to 15 million euros or 3 percent of worldwide annual turnover under Article 99; the Act’s outer ceiling of 35 million euros or 7 percent is reserved for prohibited practices such as social scoring. This blog lays out the five data decisions that belong at the board and C-suite level, what owning them actually looks like in practice, and how the major frameworks map onto them.

Key Takeaways

  • The governance gap is a data gap. Boards that engage with AI tend to govern models and use cases; the liability concentrates in data decisions about provenance, rights, quality, and regulated content, which are currently being made by default at the engineering level.
  • Five data decisions belong at the top: what data the company may train on, what data may never enter AI systems, who owns the quality metric, what flows through the vendor chain, and what evidence trail exists when a regulator or plaintiff asks.
  • Owning a decision means owning its evidence. A board that cannot see data provenance, quality metrics, and vendor attestations in its reporting pack has delegated the decision whether it intended to or not.
  • The frameworks agree more than they differ. The NIST AI Risk Management Framework, the EU AI Act, and ISO/IEC 42001 all converge on the same requirement: documented, accountable, auditable data decisions with named owners.

Why Data Decisions Are Where the Liability Lives

Think about what actually goes wrong in the AI failures that reach boards. A model trained on data the company didn’t have rights to use invites litigation that no deployment safeguard can cure. A model trained on data that underrepresents a customer population produces discriminatory outcomes that no post-hoc filter reliably catches. Customer data that entered a training set without the right consent basis creates a privacy violation that is close to irreversible, because you can’t cleanly subtract one person’s data from a trained model. In each case, the harm was locked in at the data decision, months before anyone saw an output.

The pattern is no longer hypothetical. The largest AI legal outcome to date is a training data case: the $1.5 billion copyright settlement between Anthropic and a class of book authors, granted final approval in July 2026, turned entirely on an upstream sourcing decision. The court found that training on lawfully acquired books was transformative fair use; assembling a corpus from pirate libraries was not. The US Federal Trade Commission has drawn the same line from the enforcement side, repeatedly ordering companies to delete not only improperly obtained data but the models trained on it. A provenance failure doesn’t just risk a fine. It can require destruction of the asset.

That’s why the regulatory architecture targets data directly. Article 10 of the EU AI Act requires that training, validation, and testing datasets for high-risk systems be subject to documented governance practices. Those practices cover design choices, data collection, preparation, and examination for possible biases. The commercial failure data points in the same direction. 

Gartner predicted in mid-2024 that at least 30 percent of GenAI projects would be abandoned after proof of concept by the end of 2025, listing poor data quality first among the causes. Its 2026 follow-up analysis reported the outcome was worse: at least half were abandoned after proof of concept. The legal exposure and the business-case failure share a root, and it isn’t the model.

The Five Data Decisions Boards and C-Suites Must Own

Decision 1: What Data the Company May Train On

This is the provenance and rights decision, and it’s the one with the longest liability tail. Every training dataset has a chain of custody: where it came from, under what license or consent, with what restrictions. A board doesn’t need to review datasets. It needs to know that a policy exists specifying which sourcing categories are approved (licensed, first-party with consent, commissioned collection, public domain) and which require escalation, and that someone is accountable for the provenance record on every model the company ships. In my experience, when I ask executive teams who signed off on the sourcing of their flagship model’s training data, the most common honest answer is that nobody did. It was assembled, not approved.

Decision 2: What Data May Never Enter AI Systems

The inverse decision matters as much: the categories of data that are off-limits for training, fine-tuning, or prompting regardless of business case. Health information governed by HIPAA (the US Health Insurance Portability and Accountability Act), personal data without a lawful basis under GDPR (the EU’s General Data Protection Regulation), material non-public information, privileged legal content, and customer data whose contracts exclude AI use. This boundary has to be set centrally and enforced technically, because the alternative is that it gets set implicitly by whoever is under the most delivery pressure. The test of whether this decision is owned: can management state the prohibited categories from memory, and can they show the control that enforces them?

Ownership here now extends past prevention into remediation. When prohibited data is discovered in a system after the fact, regulators have ordered deletion of the models built on it, and recent settlements have required destruction of the underlying datasets. The policy should say in advance what happens on discovery, because unwinding a trained model is expensive at best and impossible at worst.

Decision 3: Who Owns the Data Quality Metric

Data quality is the strongest single predictor of AI program failure in the published analyses, and yet in most organizations it has no executive owner: model accuracy has an owner, uptime has an owner, and the quality of the data feeding both is everyone’s job and therefore no one’s. Owning this decision means naming an accountable executive, defining the metrics (coverage, label accuracy, representativeness, freshness), and putting them in a reporting cadence that reaches the C-suite before models retrain, not after outcomes degrade. Boards should ask to see the data quality dashboard with the same expectation they’d bring to financial controls: not because directors will read every number, but because the existence and ownership of the number is the governance.

Decision 4: What Flows Through the Vendor Chain

Most enterprise AI is built on a data supply chain: annotation partners, data licensors, model providers, cloud platforms. Your compliance perimeter includes all of them. A vendor’s sourcing practices, security posture, and workforce model become your exposure the moment their output enters your training pipeline. The governance requirement is flow-down. That means contractual provenance warranties; security certifications verified rather than assumed, including ISO 27001, SOC 2, and sector-specific regimes where relevant; audit rights; and clarity about where your data physically goes and who touches it. The board-level question is simple: do we hold the same evidence about our data vendors that our customers would demand from us?

Decision 5: What Evidence Exists When Someone Asks

The last decision is about the audit trail, and it’s the one regulation has made explicit. When a regulator, plaintiff, enterprise customer, or acquirer asks how a model was trained, the answer has to exist as documentation: dataset composition, sourcing records, quality measurements, bias examinations, and the decision log of who approved what. Under the EU AI Act, this documentation is an obligation for high-risk systems; in litigation and M&A diligence, it’s rapidly becoming the default expectation for everyone else. The uncomfortable property of evidence is that it can’t be created retroactively with any credibility. The board either mandated the trail before the model shipped, or it explains the gap afterward.

What Owning These Decisions Looks Like in Practice

Ownership isn’t the board making data decisions. It’s the board ensuring the decisions have named owners, defined escalation paths, and evidence that reaches the top. In practice, that means four structures. A charter amendment placing AI data governance explicitly with a committee, typically audit or risk, so it stops being homeless on the agenda. A decision-rights matrix specifying who may approve new training data sources, who may approve exceptions to prohibited categories, and what requires escalation to the C-suite or board. A reporting pack that includes data provenance status, quality metrics, and vendor attestation status alongside the financial and cyber metrics directors already see. And a management-level review gate, so that no model ships without its data documentation complete, the same way no financial statement ships without its controls executed.

The frameworks give this structure a shared vocabulary. The NIST AI Risk Management Framework organizes it as Govern, Map, Measure, and Manage functions, with data provenance and quality sitting across all four. The EU AI Act converts the same substance into legal obligation for high-risk systems. ISO/IEC 42001, the international management-system standard for AI, packages it as an auditable management system that certification bodies can assess. A board doesn’t need to pick a winner. A practical sequence: adopt NIST as the internal organizing structure, map it to the AI Act obligations that apply to your systems, and treat ISO/IEC 42001 certification as an option when customers start asking for third-party assurance.

How Digital Divide Data Can Help

Frameworks assign the accountability; the evidence still has to be produced. Whether that layer gets built internally or with a partner, it needs to contain the same four things, and producing them is the work we do.

Provenance a regulator can read: training data with documented sourcing, licensing, and consent records, so the answer to ‘where did this data come from’ is a file rather than a reconstruction. This is what data collection and curation programs deliver.

Quality metrics an audit committee can read: measured, sampled QA with accuracy and representativeness reported continuously, the artifact Decision 3 requires an owner to produce. That reporting discipline is built into AI data preparation.

Bias examinations and evaluation evidence on a cadence: maintained, labeled evaluation sets and subgroup analyses, which is what Article 10’s examination requirement and your own board pack both draw on. Model evaluation services keep that evidence current.

And a supply chain you can flow requirements down: ISO 27001 and SOC 2 Type 2 certifications, GDPR-aligned data handling and support for HIPAA-regulated workflows where applicable, audit support, and clear answers on data residency and access, so Decision 4 holds beyond your own walls. 

If your next board pack has an AI section and it contains use cases and spend but no data provenance, quality, or vendor evidence, that’s the gap this piece is describing. Talk to an expert.

Conclusion

AI governance is arriving in boardrooms as a technology topic, and the boards that handle it well will be the ones that recognize it as a data topic, because that is where the least governed liability lives. The models will keep changing quarterly. The five decisions won’t: what we may train on, what may never enter, who owns quality, what flows through vendors, and what evidence exists when someone asks. Those decisions are being made in your organization right now, with or without governance, and the only question is whether they’re being made by the people who’ll answer for them.

The practical starting point costs one agenda item: ask management to bring the current answers to the five decisions to the next meeting, in writing, with names attached. In my experience, the value of that exercise isn’t the document. It’s the two or three blanks that nobody can fill in, because those blanks are your actual AI risk register. Delaware’s oversight doctrine gives the exercise legal weight: directors who make no good faith effort to implement reporting systems for mission-critical risks can face personal exposure, and governance counsel have begun applying that standard to AI data decisions. The blanks aren’t just a risk register. They’re the start of a defense, or the absence of one.

References

Deloitte Global Boardroom Program. (2024). Governance of AI: A critical imperative for today’s boards. https://www.deloitte.com/nz/en/services/consulting/analysis/governance-of-ai.html

Deloitte Global Boardroom Program. (2025). Progress on AI in the boardroom, but room to accelerate. https://www.deloitte.com/global/en/issues/trust/progress-on-ai-in-the-boardroom-but-room-to-accelerate.html

European Union. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European Union. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

National Institute of Standards and Technology. (2023). AI Risk Management Framework. https://www.nist.gov/itl/ai-risk-management-framework

Gartner. (2024, July 29). Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025. https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025

Gartner. (2026). Why half of GenAI projects fail: Avoid these 5 common mistakes. https://www.gartner.com/en/articles/genai-project-failure

EY Center for Board Matters. (2025). Cyber and AI oversight disclosures in 2025. https://www.ey.com/en_us/board-matters/cyber-disclosure-trends

National Association of Corporate Directors. (2025). 2025 Public Company Board Practices and Oversight Survey. https://www.nacdonline.org/all-governance/governance-resources/governance-surveys/surveys-benchmarking/2025-public-company-board-practices–oversight-survey/

The Authors Guild. (2026, July 21). Court grants final approval of $1.5 billion Anthropic copyright settlement. https://authorsguild.org/news/court-grants-final-approval-anthropic-copyright-settlement/

Mintz. (2024, January 23). Algorithmic disgorgement: An increasingly important part of the FTC’s remedial arsenal. https://www.mintz.com/insights-center/viewpoints/54731/2024-01-23-algorithmic-disgorgement-increasingly-important-part

Frequently Asked Questions

Q1. Our board isn’t technical. How can directors credibly own decisions about training data?

The same way they own financial controls without being accountants. The board’s job isn’t to evaluate datasets; it’s to verify that the decisions have named owners, documented policies, and evidence in the reporting pack. Every one of the five decisions reduces to questions a non-technical director can ask and evaluate: who approved this data source, what categories are prohibited and what enforces them, whose name is on the quality metric, what attestations do we hold from vendors, and where is the documentation. Deloitte’s finding that 79 percent of boards report limited or no AI knowledge is a case for structured questions and expert briefings, not a case for delegation by default.

Q2. We already have privacy, security, and compliance functions. Isn’t this covered?

Partially, and the gaps between the functions are exactly where AI data risk lives. Privacy governs personal data but typically has no view into whether a licensed dataset’s terms permit model training. Security governs access but not whether the data being accessed is representative or rights-cleared. Compliance tracks regulations but often maps AI obligations to no existing control owner. The five decisions are cross-functional by nature, which is why they escalate: someone with authority over all three functions has to assign the ownership, and that’s a C-suite and board-level act. A useful diagnostic is to ask each function who owns training data provenance; if you get three different answers or three referrals, it’s unowned.

Q3. Which framework should we adopt: NIST AI RMF, ISO/IEC 42001, or the EU AI Act?

They’re not competitors, and the practical answer is a sequence rather than a selection. The EU AI Act isn’t optional if your systems fall in its scope; it’s law, and its data governance article defines obligations, not suggestions. The NIST AI Risk Management Framework is voluntary and works well as the internal organizing structure because it’s function-based and framework-agnostic. ISO/IEC 42001 matters when you need third-party assurance, because it’s the one a certification body can audit against, and enterprise customers are beginning to ask for it in procurement the way they ask for ISO 27001 today. The pattern most organizations land on: NIST for structure, the AI Act for legal floor, and 42001 certification when the market demands the certificate.

Q4. What should actually appear in the board reporting pack for AI data governance?

Five artifacts, one per decision, each fitting on a page. A provenance summary: models in production, data sources per model, approval status, and any sources under remediation. A prohibited-data attestation: the categories, the enforcing controls, and any exceptions granted with their approvers. The quality dashboard: the owned metrics with trend lines and threshold breaches. A vendor status table: data supply chain partners, certifications verified, attestations current or expired. And a documentation readiness indicator: which production models have complete data documentation and which have gaps. The pack’s purpose isn’t detail; it’s that a director can see in five pages whether the five decisions are owned and evidenced, and can ask about anything red.

Q5. We don’t operate in Europe. Does the EU AI Act really matter to us?

Quite possibly, and the determination belongs with counsel rather than a blog, but two facts are worth knowing before that conversation. The Act’s reach extends beyond companies established in the EU: providers placing systems on the EU market and situations where system outputs are used in the EU can fall in scope regardless of where the company sits. And even for companies genuinely outside its reach, the Act is functioning as the reference standard: enterprise customers, investors, and other regulators are borrowing its categories and its documentation expectations, which means its data governance requirements describe the evidence sophisticated counterparties will ask for irrespective of jurisdiction. Building the documentation trail only for the markets that legally require it usually costs more than building it once.

AI Governance Frameworks: What Boards and C-Suites Need to Own About Data Decisions Read Post »

AI-powered warehouse robots and a monitoring dashboard representing automated data curation and dataset quality management.

What AI Data Curation Really Involves Beyond Data Cleaning

AI data curation is the active, ongoing practice of deciding what belongs in a training dataset, in what proportion, with what documented origin, and with what evidence that the mix matches the task the model will perform. Data cleaning removes errors from records that are already in hand. Curation determines which records should be in hand at all, which means a dataset can be completely clean and still be the wrong dataset.

The distinction matters because teams keep spending their quality control budget in the wrong place. Deduplication scripts, null-value handling, and format normalization are cheap to run and easy to measure, so they get done. Coverage planning, diversity scoring, and provenance tracking are harder to measure, so they get deferred until a model underperforms in production and nobody can explain why. Data collection and curation services address the selection layer that sits above cleaning, while Data preparation services handle the transformation and structuring work that follows selection.

Key Takeaways

  • Data cleaning fixes errors in the records you already have, while data curation decides which records should be in the dataset at all.
  • A dataset can pass every cleaning check and still be the wrong dataset, which is why curation failures usually show up only after a model reaches real users.
  • Recent open research shows that better-chosen training data beats simply adding more of it, and it cuts the cost of training at the same time.
  • The most reliable way to curate is to write down what the finished dataset should look like before collecting anything, then measure what you actually gathered against that target.
  • Problems like uneven representation are invisible to error-checking tools, because unbalanced records are not broken records.
  • Recording where every piece of data came from has to happen while the dataset is being built, since that history cannot be recreated later.

What is AI data curation and where does it sit in the AI pipeline?

AI data curation is the deliberate selection, organization, enrichment, and maintenance of data so that a dataset is fit for training a specific model against a specific objective. It sits between raw data acquisition and model training, and it stays active after deployment as the target distribution shifts. Curation covers source selection, coverage planning, filtering, deduplication, labeling design, metadata capture, and lineage documentation. Data annotation solutions are one component inside that scope rather than a substitute for it.

Terminology in this area is inconsistent across vendors, so it helps to fix definitions before going further. Data curation, dataset curation, and training data curation refer to the same practice at different levels of specificity. Data cleaning, sometimes written as data cleansing, is a subset of curation concerned with correcting errors in records that already exist. Data governance covers the policies that constrain how data may be acquired, stored, and used. Data management covers the infrastructure that stores and serves it. Curation is the editorial function that runs on top of all three.

What is the difference between data curation and data cleaning?

The practical test for whether a team is curating or cleaning is simple. Cleaning asks whether each record is correct. While, curation asks whether the collection, taken as a whole, teaches the model the distribution it will encounter. 

Data cleaning is corrective and bounded. It operates on a dataset that has already been assembled, and its success criterion is the absence of defects: no malformed timestamps, no duplicate rows, no impossible values, no missing required fields. The work is largely rule-driven, it can be automated to a high degree, and it terminates. Once the defect rate falls below threshold, cleaning is finished until new data arrives.

Data curation is compositional and open-ended. It operates on the question of what the dataset should contain, which means it involves judgments that no rule can settle on its own: how much of each domain, which edge cases deserve overrepresentation, which sources to exclude on licensing grounds, which annotator populations to recruit for which categories. The work is partly automated and partly human, and it does not terminate, because the deployment environment keeps moving. Building AI-ready datasets requires a clear understanding of where these decisions occur across the data pipeline and the failure modes that can emerge at each stage.

The two practices differ across four dimensions that matter for planning and budgeting:

Dimension Data cleaning Data curation
Unit of analysis The individual record The dataset as a distribution
Core question Is this value correct? Should this example be here, and in what proportion?
Failure signature Training crashes, obvious label noise, schema errors Model performs well on benchmarks and fails on production traffic
Endpoint Terminates when defect rate clears threshold Continuous; re-run as deployment distribution shifts

The failure signature row is the one worth dwelling on. Cleaning failures are loud, because broken records tend to break pipelines. Curation failures are quiet. A narrow dataset produces a model that scores well on a held-out split drawn from the same narrow distribution, then degrades on the traffic that matters. By the time the gap appears, the training run is months old and the diagnosis is expensive.

Why is data curation important for AI model quality?

The empirical case for curation has strengthened considerably since 2024, largely because open dataset research made controlled comparisons possible for the first time. The FineWeb dataset study documented and ablated each filtering and deduplication decision applied to 96 Common Crawl snapshots, and showed that the curation recipe itself, rather than corpus size alone, drove downstream benchmark performance. Its educational subset, filtered from the same underlying pool, produced markedly stronger results on knowledge and reasoning benchmarks.

The DataComp-LM benchmark made the same point under controlled conditions across model scales from 412M to 7B parameters. Holding architecture and training recipe fixed and varying only the curation strategy, the study found that model-based filtering was the decisive factor in assembling a high-quality training set, and that a better-curated corpus reached higher accuracy with substantially fewer training tokens. Curation converts directly into compute savings, which is the argument that tends to land with budget holders.

Generative systems amplify the effect because their outputs are open-ended. A classifier trained on a skewed dataset produces measurable error on the underrepresented class. A generative model trained on the same skew produces fluent, confident output that reflects the skew without flagging it. Hallucinations, fine-tuning instability, and representational bias often originate in data composition decisions made long before model training begins.

How do you curate a training dataset step by step?

Curation becomes tractable when it is treated as a sequence with defined artifacts at each stage. The sequence below reflects how mature programs structure the work. The order matters, because steps taken out of sequence produce datasets that are internally consistent and externally wrong.

  1. Write the target specification first: Define what the finished dataset should look like before collecting anything: domains, languages, modalities, edge-case categories, minimum counts per stratum, and acceptance thresholds. Teams that skip this step end up with whatever was easiest to acquire, and they discover the shape of their dataset only after training.
  2. Map sources against the specification: Identify which sources can supply which strata, and record the gaps explicitly. Gaps that are known in advance can be filled through targeted collection or synthetic augmentation. Gaps discovered after training cannot.
  3. Filter for relevance before filtering for quality: Relevance filtering removes material that is well-formed and irrelevant to the task. Quality filtering removes material that is relevant and defective. Running quality filters first wastes effort on records that will be discarded anyway.
  4. Deduplicate at three levels: Exact duplicates are trivial to remove. whereas Near-duplicates require fuzzy matching such as MinHash, and Semantic duplicates require embedding-based similarity. Aggressive thresholds reduce redundancy and also strip legitimate variation, so the threshold is a tuning decision rather than a default.
  5. Score diversity and coverage against the specification: Measure the assembled dataset against the strata defined in step one and report the deltas. Coverage reporting is the artifact that distinguishes a curated dataset from a large one.
  6. Annotate with iterative guideline development: Labeling schemas rarely survive first contact with real data. Run pilot batches, measure inter-annotator agreement, revise the guidelines, and re-run. Agreement scores are the instrument that tells you whether the schema is well-defined.
  7. Validate, document, and schedule the next cycle: Produce a datasheet recording sources, licenses, transformations, exclusions, and known limitations. Then set the review interval, because the deployment distribution will move.

Synthetic data has a defined role within this sequence. It is most valuable for addressing known coverage gaps, particularly in rare-event scenarios and privacy-constrained domains. However, it should complement rather than replace human-curated data, as synthetic generation can introduce artifacts, unrealistic patterns, and hidden distortions that rigorous validation must identify before the data is used for training.

How does curation surface bias that cleaning leaves untouched?

Cleaning cannot detect representational bias, because biased records are not defective records. A facial recognition corpus in which 85 percent of images depict light-skinned subjects contains no malformed files, no missing fields, and no label errors. Every cleaning check passes. The dataset is nonetheless unusable for deployment across a general population, and the only stage at which the problem is visible is the stage that measures composition against a target.

Bias enters datasets through several distinct channels, and each requires a different curation control. Measurement bias comes from instruments that distort systematically, such as miscalibrated sensors or low-fidelity audio capture. Sample bias comes from source populations that do not match the deployment population. Cultural and linguistic bias comes from annotator populations whose conventions differ from those of end users. Data bias in AI training sets works through concrete cases in each category, including how regional vocabulary differences in annotation teams produce systematically wrong labels.

Three curation controls address these channels directly:

  • Stratified coverage audits that compare dataset composition against the demographic and contextual profile of the deployment environment, run before training rather than after evaluation.
  • Annotator population design that matches the linguistic and cultural context of the target users, with agreement measured separately across annotator groups to expose systematic divergence.
  • Data-level correction through resampling, reweighting, or targeted collection, applied to the dataset rather than compensated for through post-hoc model adjustments that are harder to document and audit.

Why does provenance tracking belong inside curation, not compliance?

Provenance is frequently treated as a legal formality handled after the dataset is built. That sequencing fails, because lineage that was not captured during assembly cannot be reconstructed afterward. The Data Provenance Initiative audit traced over 1,800 widely used text datasets and found license omission rates above 70 percent and license error rates above 50 percent on popular hosting sites. Teams building on public corpora are frequently operating with incorrect information about what they are permitted to use.

Provenance also has an engineering function that has nothing to do with licensing. When a model exhibits a specific failure mode, the diagnostic question is which subset of training data produced it. Answering that requires per-record lineage: source, acquisition date, transformation history, annotation batch, and reviewer. Programs that capture this during curation can isolate and correct the responsible subset. Programs that did not capture it retrain from scratch and hope. Structured metadata makes lineage capture a routine part of dataset assembly.

Regulatory pressure is converging on the same requirement. Documentation obligations for training data are becoming a condition of deployment in several jurisdictions, and the datasheet produced in step seven of the curation sequence is the artifact that satisfies them. Programs that already produce it for engineering reasons absorb the compliance requirement at close to zero marginal cost.

What tools help with AI dataset curation?

No single tool covers curation end to end, and treating any one of them as a complete solution is a common and expensive mistake. The tooling landscape divides into functional categories, and a working stack draws from several.

  • Deduplication and filtering frameworks: MinHash and SimHash implementations for near-duplicate detection, embedding-based semantic deduplication, and model-based quality classifiers of the kind used in FineWeb and DataComp-LM. These handle volume, and they encode the thresholds that determine dataset diversity.
  • Dataset exploration and curation platforms: Tools that support visual inspection, embedding-space clustering, similarity search, and slice-based analysis of large image and video corpora. Their value is in making distribution gaps visible to a human reviewer.
  • Label quality and error detection: Confident-learning libraries and agreement-analysis tooling that surface probable label errors and annotator drift, which manual review misses at scale.
  • Lineage, versioning, and documentation: Dataset versioning systems and metadata catalogs that make datasets reproducible and auditable, so that a training run can be tied back to an exact dataset state.
  • Annotation platforms with quality instrumentation: Systems that support iterative guideline revision, multi-pass review, and inter-annotator agreement reporting as first-class features rather than exports.

The judgment layer stays human regardless of tooling. Tools measure duplication rates, agreement scores, and embedding density. Deciding which coverage gap matters most for a given deployment, which edge cases justify overrepresentation, and where a diversity threshold should sit remains a design decision informed by domain knowledge.

Where do AI data curation services fail in practice?

Curation programs tend to fail in four recognizable ways, and all four are structural rather than technical. Naming them is useful, because each has a specific organizational remedy.

  • Curation is scoped as a one-time project: A team curates a dataset, ships a model, and moves on. Within a year the deployment distribution has shifted and dataset quality has effectively degraded, even though no file changed. The remedy is a scheduled review cycle tied to model retraining.
  • Cleaning metrics are used as curation metrics: Defect rates and completeness percentages are reported as evidence of dataset quality. They measure hygiene and say nothing about coverage. The remedy is to report composition against the target specification alongside defect rates.
  • Curation runs only downstream: Effort concentrates on correcting problems in data that has already been collected, when the cheapest intervention point is the collection design itself. The remedy is to move specification and source mapping ahead of acquisition.
  • Over-curation narrows the dataset: Aggressive filtering and deduplication remove noise and also remove the legitimate variation that produces robustness. The remedy is to treat every filtering threshold as a tuned parameter, validated against held-out performance rather than set by default.

How Digital Divide Data Can Help

DDD operates curation as a full pipeline function rather than a labeling engagement. Data collection and curation services cover source identification and coverage planning at the front of the pipeline, deduplication and quality filtering in the middle, and post-curation validation against the target specification at the end. Diversity planning is structured across languages, domains, demographic groups, and content types, so that dataset assembly targets the coverage gaps that affect model behavior rather than the dimensions that are simplest to source at volume.

On the quality side, annotation programs run with iterative guideline development, multi-pass review, and inter-annotator agreement measured per category and per annotator cohort, which is how systematic divergence between annotator groups becomes visible before it reaches the training set. Trust and safety solutions extend this into bias and fairness auditing, applying stratified composition audits and data-level correction before training rather than post-hoc adjustment afterward. DDD’s global delivery footprint supports annotator populations matched to the linguistic and cultural context of the deployment environment, including low-resource languages where representative data is hardest to source.

Lineage is captured during assembly. Source, acquisition date, transformation history, annotation batch, and reviewer are recorded per record, which produces the datasheet needed for regulatory documentation and the diagnostic trail needed to isolate a problematic subset when a model misbehaves in production.

Build training datasets that hold up in production, not just in evaluation. Talk to an Expert

Conclusion

Cleaning answers whether the records in hand are correct. Curation answers whether those are the right records, in the right proportions, from documented sources, measured against the distribution the model will actually meet. The second question is harder to instrument and it is the one that determines whether a model survives contact with production traffic.

Organizations that treat curation as an ongoing editorial discipline accumulate an asset: a dataset with known composition, documented lineage, and a review cadence that keeps it aligned as conditions change. Organizations that treat it as pre-processing accumulate a liability that stays invisible until a model underperforms and nobody can trace why. The gap between the two compounds with every retraining cycle. 

References

Penedo, G., Kydlíček, H., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L., & Wolf, T. (2024). The FineWeb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems. https://arxiv.org/abs/2406.17557

Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S., Bansal, H., Guha, E., Keh, S., Arora, K., Garg, S., Xin, R., Muennighoff, N., Heckel, R., Mercat, J., Chen, M., Gururangan, S., Wortsman, M., Albalak, A., Bitton, Y., Nezhurina, M., Abbas, A., Hsieh, C.-Y., Ghosh, D., Gardner, J., Kilian, M., Zhang, H., Shao, R., Pratt, S., Sanyal, S., Ilharco, G., Daras, G., Marathe, K., Gokaslan, A., Zhang, J., Chandu, K., Nguyen, T., Vasiljevic, I., Kakade, S., Song, S., Sanghavi, S., Faghri, F., Oh, S., Zettlemoyer, L., Lo, K., El-Nouby, A., Pouransari, H., Toshev, A., Wang, S., Groeneveld, D., Soldaini, L., Koh, P. W., Jitsev, J., Kollar, T., Dimakis, A. G., Carmon, Y., Dave, A., Schmidt, L., & Shankar, V. (2024). DataComp-LM: In search of the next generation of training sets for language models. arXiv preprint. https://arxiv.org/abs/2406.11794

Longpre, S., Mahari, R., Chen, A., Obeng-Marnu, N., Sileo, D., Brannon, W., Muennighoff, N., Khazam, N., Kabbara, J., Perisetla, K., Wu, X., Shippole, E., Bollacker, K., Wu, T., Villa, L., Pentland, S., & Hooker, S. (2023). The Data Provenance Initiative: A large scale audit of dataset licensing and attribution in AI. arXiv preprint. Published in Nature Machine Intelligence (2024). https://arxiv.org/abs/2310.16787

Frequently Asked Questions

What is AI data curation in simple terms?

It is the work of deciding what goes into a training dataset and keeping those decisions documented and current. That covers choosing sources, setting how much of each type of data you need, filtering what does not belong, labeling what remains, and recording where everything came from.

Is data cleaning part of data curation, or a separate thing?

Cleaning is one step inside the curation sequence. Cleaning fixes errors in records you already have. Curation decides which records you should have in the first place, which is a broader job that keeps running after the cleaning is done.

Can a dataset be perfectly clean and still be bad for training?

Yes, and this is the most common way training data fails. A dataset with no formatting errors, no duplicates, and no missing fields can still cover only a narrow slice of what the model will meet in production. Every cleaning check passes and the model still fails on real traffic.

How often should a training dataset be re-curated?

Tie the review to your retraining schedule rather than to a fixed calendar. The environment a model operates in keeps shifting, so a dataset that matched it a year ago may no longer match it now, even though not a single file has changed.

What AI Data Curation Really Involves Beyond Data Cleaning Read Post »

Stages of AI Data Preparation for Production-Ready Training Data

The 7 Stages of AI Data Preparation for Production-Ready Training Data

AI data preparation services convert raw, inconsistent source data into training-ready datasets through seven stages: raw intake, deduplication, normalization, format conversion, augmentation, quality scoring, and export/delivery. Most teams underinvest in deduplication and quality scoring, which is where duplicate contamination and undetected label noise enter the training set. A full preparation cycle typically runs two to twelve weeks, depending on volume, modality, and whether the source data arrived with usable provenance.

The decision facing most ML engineering teams is not whether to prepare data. It is whether to build the pipeline in-house or source it. AI data preparation services exist because the work is high-volume, judgment-heavy, and unglamorous, and because doing it repeatedly costs more infrastructure than most teams budget for. The same provenance tracking and sampling logic governs data collection and curation services, which is why the two functions are usually brought together. 

Key Takeaways

  • Data preparation is the work of turning raw, messy source data into a clean dataset a model can actually learn from, and it runs across seven stages: intake, removing duplicates, standardizing, converting formats, filling coverage gaps, scoring quality, and exporting.
  • Most projects fail on the data rather than the model, so the effort spent here is what separates systems that work in production from ones that only look good in testing.
  • Removing duplicates is the step teams most often rush, even though repeated content wastes labeling budget and quietly inflates test scores.
  • Preparation comes before labeling, and reversing that order means paying people to label content you were going to throw away.
  • A finished dataset should arrive with a record of where every piece came from, what was done to it, and a split that keeps the same source out of both training and testing.
  • Timelines swing from a couple of weeks to a few months depending mostly on how much you already know about where your data came from.

What is AI data preparation, and where does it sit in the ML lifecycle?

AI data preparation is the sequence of transformations that turns raw source data into a dataset a model can train on. It sits between data collection and model training, and covers ingestion, deduplication, cleaning, standardization, encoding, and validation. Practitioners also call it data preprocessing, data wrangling, or AI data prep; the terms are interchangeable. Data engineering for AI supplies the infrastructure underneath including orchestration, storage, and lineage tracking, that lets these transformations run as a pipeline rather than as one-off notebooks.

The distinction that matters most to engineering teams is between preparation and labeling. Preparation operates on the data itself: its structure, format, distribution, and integrity. Labeling operates on the meaning attached to it. A dataset can be perfectly labeled and still be unusable if it contains near-duplicates, inconsistent units, or a split that leaks. DDD’s earlier work on ML data preparation made a point that has aged well: preparation consumes most of a data team’s time precisely because it is the most consequential part of the job.

Data preparation is also where most production failures begin. RAND’s interview study of 65 data scientists and engineers reported that more than 80 percent of AI projects fail, roughly twice the failure rate of IT projects without AI, with inadequate data among the five leading root causes. Model architecture is rarely the binding constraint. The dataset is.

What are the seven stages of a production AI data preparation workflow?

Each stage below produces two things: a transformed artifact and a check that the transformation did what it was supposed to. Skipping the check is how teams end up with pipelines that run cleanly and produce datasets nobody can trust.

Stage 1: Raw Data Intake to establish Provenance

Intake is where a dataset acquires its audit trail. Every incoming file, record, or sensor sequence gets registered with a source identifier, a timestamp, a license or consent basis, and a checksum. Teams that skip this cannot later answer basic questions: where did this record come from, were we permitted to use it, and has it changed since ingestion. Intake also fixes the sampling frame, which determines whether coverage gaps are even visible later on.

Three artifacts are worth producing at this stage:

  • A source registry: One row per source, recording license, consent basis, and collection date.
  • Checksums on ingest, so silent corruption is detectable rather than mysterious.
  • A coverage baseline recording what the dataset contains along the dimensions you care about, including geography, language, lighting condition, demographic slice, and vehicle class.

Stage 2: Deduplication

Duplicates inflate the apparent size of a dataset while shrinking its effective information content. In text corpora, near-duplicate documents drive memorization and quietly contaminate benchmarks when the same passage lands in both training and evaluation splits. The FineWeb dataset ablations showed that deduplication strategy measurably changed downstream model performance across a 15-trillion-token corpus.

Exact-match deduplication is cheap and catches very little. Production pipelines run three passes:

  • Exact hashing on raw bytes or normalized text, which removes the trivial cases.
  • Fuzzy matching: MinHash with locality-sensitive hashing for text, perceptual hashing for images to catch near-duplicates that differ by formatting or compression.
  • Semantic deduplication using embeddings, which catches records that convey the same content in different surface forms.

In perception and ADAS datasets, the equivalent problem is temporal redundancy. Consecutive frames from a stationary vehicle are nearly identical and add annotation cost without adding signal. DDD’s guide to building datasets for large language model fine-tuning works through the text-side version of the same trade-off in more depth.

Stage 3: Normalization to standardize

Normalization strips out variation that carries no signal. In tabular data that means units, encodings, date formats, null representations, and categorical vocabularies. In text it means Unicode normalization, casing and whitespace rules, and consistent handling of boilerplate. In sensor data it means coordinate frames, timestamp alignment across cameras and LiDAR, and calibration metadata.

Sensor synchronization deserves particular attention. A multi-organization study of annotation quality across six automotive companies found that synchronization and calibration issues were a recurring completeness error; unsynchronized sensors produce annotations that drift in space and time, which degrades multimodal fusion downstream. Normalization is where that gets caught, before anyone spends money labeling misaligned frames.

Semantic normalization is the harder half of the stage. Acronyms, jargon, and domain vocabularies have to resolve to consistent entities. 

Stage 4: Format Conversion for the Training Loop

Format conversion turns normalized records into the physical layout the training job actually reads. In practice, that means columnar or sharded formats; Parquet, Arrow, WebDataset, or TFRecord, sized so each shard streams to the accelerators without starving them. For multimodal data it also means deciding what lives inline in the shard and what lives as a pointer to object storage.

Three decisions at this stage have long consequences:

  • Shard size and count, which govern shuffle quality and read throughput.
  • Schema versioning, so a dataset regenerated in six months is still readable by the training code that consumed the original.
  • Tokenization and encoding boundaries, which are effectively irreversible once baked into the shards.

Stage 5: Data Augmentation

Augmentation expands coverage where real data is scarce. For vision, that means geometric and photometric transforms, synthetic weather and lighting, and simulated rare events. For text, it means paraphrase, back-translation, and instruction reformatting. The purpose is to harden the model against variation it will meet in production and has not seen enough of in training.

Augmentation stops helping when it starts distorting the distribution. Two failure modes recur: augmenting the majority class and widening an imbalance that was already there, and training recursively on synthetic outputs until diversity collapses. The rule that survives contact with production is to augment against a measured coverage gap. If the Stage 1 coverage baseline does not show a gap, augmentation is adding cost without adding capability.

Stage 6: How do you score dataset quality before training?

Quality scoring assigns a measurable value to records and to the dataset as a whole, so filtering decisions are defensible rather than intuitive. It operates on three levels. Record-level scoring flags corrupt files, truncated sequences, low-information samples, and out-of-distribution records. Label-level scoring measures inter-annotator agreement, isolates disagreement clusters, and surfaces suspected label errors. Dataset-level scoring measures class balance, coverage against the sampling frame, and drift against the production distribution.

The automotive study cited above catalogued 18 recurring annotation error types across three dimensions: completeness, accuracy, and consistency, and the practitioners who reviewed it described the result as a failure-mode catalogue comparable to FMEA. That is the right mental model for this stage. Quality scoring is a diagnostic that tells you which errors you have and how many, not a pass/fail gate. Data quality defines the success of AI systems from the model-behavior side.

Stage 7: Leakage-safe Export

Export is where the dataset becomes an immutable, versioned artifact. Three things have to be true. The split must be leakage-safe; records sharing an entity, a session, or a source document belong in the same split, or evaluation metrics will be optimistic and will not reproduce in production. The dataset must be versioned, with a manifest recording every transformation applied. And it must carry its documentation, usually a datasheet describing sources, consent basis, known gaps, and intended use, which is also what emerging AI regulation increasingly expects for high-risk systems.

Leakage is the quietest failure in the entire workflow. It produces no error and breaks no job. It produces a model that looks better than it is, and the gap only reveals itself after deployment.

What is the difference between data preparation and data annotation?

Preparation and annotation are sequential steps, not alternatives. Preparation acts on the data; annotation adds meaning to it. Deduplicating a corpus, aligning LiDAR timestamps, and converting to Parquet are preparation. Drawing a 3D cuboid around a pedestrian or tagging a support ticket as billing-related is annotation, and it belongs to multimodal data annotation services.

The order has direct cost consequences. Annotating a corpus before deduplicating it means paying to label the same content more than once. Annotating sensor data before validating calibration means labeling frames that will later be discarded. Teams that treat preparation as a prerequisite for annotation consistently spend less than teams that treat it as cleanup afterwards.

What tools are used for AI data preparation, and how long does it take?

No single tool covers the workflow. A production stack usually combines:

Orchestration: Airflow, Dagster, or Prefect, to schedule stages and retry failures.

Distributed processing: Spark, Ray, or Dask, once volume exceeds a single machine.

Deduplication: MinHash/LSH libraries for text, perceptual hashing for images, embedding-based semantic dedup for the hard cases.

Validation: Great Expectations, Deequ, or Pandera, for schema and distribution assertions.

Dataset versioning: DVC, LakeFS, or Delta Lake, so an artifact can be regenerated exactly.

Curation and visual QA: FiftyOne or equivalent, for image, video, and point-cloud inspection.

Timelines depend on three variables: volume, modality, and the quality of the provenance that arrived with the data. A structured tabular dataset with clean lineage can move through all seven stages in two to three weeks. A multimodal corpus assembled from heterogeneous sources with no source registry more often takes eight to twelve weeks, and a disproportionate share of that goes to Stage 1, because provenance has to be reconstructed rather than simply recorded. Sensor datasets sit in between and are usually dominated by calibration and synchronization work.

When should AI data preparation be sourced as a managed service?

Building the pipeline in-house is the right call when the data is highly proprietary, the transformations are stable, and the team already employs data engineers who are not otherwise committed. That combination is rarer than it appears. The recurring reason teams outsource is not a capability gap. 

Five questions separate a serious preparation partner from a reseller:

  • Do they deduplicate beyond exact match, and can they show you the pass structure?
  • Do they deliver a dataset manifest and datasheet, or just a folder of files?
  • Can they demonstrate leakage-safe splitting on entity-grouped or session-grouped data?
  • Are they toolchain-agnostic, or is everything routed through one platform they happen to resell?
  • Do their security certifications actually cover the data class you are handing over?

How Digital Divide Data Can Help

DDD runs the full preparation lifecycle as a managed program. Intake, deduplication, normalization, format conversion, augmentation, quality scoring, and export are delivered through our data pipeline services, with human-in-the-loop review concentrated at the stages where automation is least reliable: semantic normalization, label-error adjudication, and coverage assessment against the sampling frame. Our teams are toolchain-agnostic and work inside the client’s existing stack rather than migrating data into a proprietary platform.

For Physical AI, ADAS, and autonomous vehicle programs, preparation is dominated by multi-sensor alignment. Our sensor data annotation teams handle timestamp synchronization, calibration validation, and cross-modality projection checks before any labeling begins, which is where a large share of downstream perception error is prevented rather than corrected later. For generative AI programs, the same discipline applies to corpus deduplication, contamination screening against evaluation benchmarks, and provenance documentation. Delivery operates under SOC 2 Type 2 and ISO 27001 controls, with GDPR and HIPAA handling where the data class requires it.

Move your training data from raw intake to a versioned, leakage-safe artifact with Digital Divide Data.

Conclusion

The seven stages are not a checklist to run once before the interesting work begins. They are a loop that runs every time the data changes, and the organizations that treat them that way end up with datasets they can audit, reproduce, and improve. The organizations that treat preparation as a one-time cleanup tend to find their problems in production, where a fix costs an order of magnitude more than it would have cost upstream.

That gap is widening. As models become cheaper to train and easier to swap, the dataset becomes the durable asset. Teams that can regenerate a dataset from a manifest, explain every filter they applied, and prove their splits are clean will move faster — not because their models are better, but because they can trust their own numbers. 

References

Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L., & Wolf, T. (2024). The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. https://arxiv.org/abs/2406.17557

Ryseff, J., De Bruhl, B. F., & Newberry, S. J. (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI. RAND Corporation, Report RR-A2680-1. https://www.rand.org/pubs/research_reports/RRA2680-1.html

Saeeda, H., Johansson, T., Mohamad, M., & Knauss, E. (2025). Data Annotation Quality Problems in AI-Enabled Perception System Development. arXiv preprint arXiv:2511.16410. https://arxiv.org/abs/2511.16410

Frequently Asked Questions

What is AI data preparation?

It is the work of turning raw source data into a dataset a model can actually train on. That covers taking the data in, removing duplicates, standardizing formats and units, converting it into a training-ready file layout, scoring its quality, and exporting a versioned copy with clean train/test splits.

How long does AI data preparation take?

It depends on volume, data type, and how much you know about where the data came from. Clean tabular data with good records can be through the whole workflow in two to three weeks. A messy multimodal collection with no source history usually takes eight to twelve weeks, mostly because someone has to reconstruct the provenance before anything else can start.

What is the difference between data preparation and data annotation?

Preparation changes the data by deduplicating it, aligning sensor timestamps, converting file formats. Annotation adds meaning to it, like drawing a box around a pedestrian or tagging a comment as a complaint. Preparation comes first, and doing it in that order saves money, because you are not paying to label content you would have thrown away anyway.

What tools are used for AI data preparation?

There is no single tool. Most teams stitch together an orchestrator like Airflow or Dagster, a distributed engine like Spark or Ray, deduplication libraries such as MinHash or perceptual hashing, a validation layer like Great Expectations, and a versioning system like DVC or Delta Lake. Visual QA tools such as FiftyOne cover image, video, and point-cloud review.

The 7 Stages of AI Data Preparation for Production-Ready Training Data Read Post »

Audit an AI Model for Bias

How to Audit an AI Model for Bias: A Practical Data-Level Checklist

Kevin Sahotsky

Bias in AI models is overwhelmingly a data problem before it is a model problem. The patterns a model learns, the groups it overrepresents or underrepresents, and the shortcuts it takes when making predictions. Almost all of these trace back to characteristics of the data the model was trained on. This is particularly relevant for AI program leads, product managers overseeing model deployments, and compliance teams working in regulated industries where demonstrating fairness is not optional.

This blog walks through a practical data-level checklist for auditing an AI model for bias, covering where bias enters, what to measure, and what the remediation options actually look like. Trust and safety solutions and model evaluation services are the two capabilities most directly involved in identifying and addressing data-level bias before it reaches production.

Key Takeaways

  • Bias in AI models originates in training data far more often than in model architecture. Auditing the architecture without auditing the data misses the root cause.
  • There are three stages where bias enters: data collection, data labeling, and data curation. Each stage requires its own audit approach and cannot be substituted by checks at the other stages.
  • Representation gaps are the most common and most overlooked source of bias. A model trained on data that systematically underrepresents certain groups will produce worse outputs for those groups even when no individual annotation is wrong.
  • Fairness metrics measure different things and can contradict each other. Choosing which metric to optimize requires an explicit decision about what kind of fairness matters for the deployment context.

Where Bias Actually Comes From

Stage 1: Data Collection

The first place bias enters is at collection. If the data collected to train a model does not represent the full range of people, contexts, and conditions the model will encounter at deployment, the model will systematically underperform on the cases that were underrepresented in training. This is not a labeling problem. The labels can all be correct, and the model will still produce biased outputs because it has seen too few examples of certain groups or conditions to learn to handle them well.

Collection bias is the hardest to fix after the fact because it requires going back and collecting more data from the underrepresented cases, which is expensive and time-consuming. The audit question at this stage is simple but easy to defer: does the distribution of the training data match the distribution of the deployment population? Data collection and curation services that audit demographic and contextual coverage before collection ends are far cheaper than auditing after a biased model has reached production.

Stage 2: Data Labeling

The second entry point is labeling. Human annotators apply labels to training data, and those labels reflect the annotators’ own frames of reference, cultural contexts, and implicit associations. An annotator who consistently associates certain names with certain characteristics, or who applies sentiment labels differently across different dialects or writing styles, introduces label-level bias that the model will learn directly. Because label bias looks like signal rather than noise from the model’s perspective, it is often harder to detect than representation gaps.

The audit approach at this stage is inter-annotator agreement disaggregated by subgroup. If annotators agree consistently on majority-group examples but diverge significantly on minority-group examples, the annotation process is introducing differential error rates that the model will inherit. Text annotation services that measure inter-annotator agreement at the subgroup level, not just in aggregate, surface this pattern before it compounds through the full training dataset.

Stage 3: Data Curation

The third entry point is curation. Even when collection and labeling are unbiased, the decisions made about which data to keep, which to filter, and how to balance the training set introduce bias. A curation pipeline that filters out low-confidence examples disproportionately removes data from underrepresented groups, because low-confidence labeling correlates with the annotators’ lower familiarity with those groups. A resampling strategy that balances by category but not by demographic subgroup within category can leave systematic gaps.

Curation bias is the most invisible of the three because it happens in the pipeline rather than in the data itself. The audit requires tracking not just what data was kept but what was removed and why, which most curation pipelines do not do by default.

The Data-Level Bias Audit Checklist

Check 1: Representation Audit

Map the demographic and contextual distribution of your training data against the deployment population. For each group that matters for your deployment context, calculate the proportion in the training set versus the proportion in the population the model will serve. A gap of more than ten percentage points between a group’s representation in training and its representation in the deployment population is a useful starting threshold for flagging meaningful risk, warranting either additional data collection or a fairness constraint during training. The right threshold will vary with deployment context and the stakes involved.

Representation audit tools include demographic classifiers applied to the training set, metadata analysis where demographic fields exist, and external benchmarks that characterize the expected deployment distribution. The output is a coverage map, not a single metric.

Check 2: Label Consistency Audit

Calculate inter-annotator agreement disaggregated by the subgroups relevant to your deployment context. The relevant breakdown depends on the application: for a hiring model, this might be by applicant name type or inferred demographic; for a content moderation model, this might be by dialect or topic type; for a medical model, this might be by patient demographic characteristics in the case descriptions.

As a useful starting threshold, any subgroup showing inter-annotator agreement more than ten percentage points below the overall agreement level is a signal worth investigating, suggesting the labeling process may be applying different standards to different groups. This is the input to annotator calibration and guideline revision, not a reason to discard the data. Model evaluation services that measure subgroup-level annotation consistency as a standard output of the labeling quality process catch this before it accumulates through the full training set.

Check 3: Curation Audit

Document what was removed from the training set and why. For each filtering step, calculate the removal rate disaggregated by subgroup. If a low-confidence filter removes data from one subgroup at twice the rate of another, that filter is introducing a representation gap that did not exist in the raw collected data. The audit does not require abandoning confidence-based filtering. It requires checking whether the filter is applied uniformly across groups and adjusting the threshold or supplementing with additional collection where it is not.

Check 4: Performance Disparity Measurement

Evaluate model performance disaggregated by subgroup across your held-out evaluation set. The relevant metrics depend on the task. For classification tasks, measure precision, recall, and F1 separately for each subgroup. For regression tasks, measure mean error and error variance. For generative tasks, use human evaluation panels drawn from the relevant subgroups rather than automated metrics, because automated metrics often have their own demographic biases.

Performance disparity greater than five percentage points in recall across demographic subgroups on a classification task in a regulated domain is a reasonable benchmark for a material finding requiring remediation before deployment, though the appropriate threshold depends on the regulatory context and the consequences of false negatives for each subgroup.

Check 5: Fairness Metric Selection

Different fairness metrics operationalize different concepts of fairness, and they can mathematically conflict with each other. Demographic parity requires that the positive prediction rate is equal across groups. Equalized odds requires that both the true positive rate and the false positive rate are equal across groups. Calibration requires that predicted probabilities correspond to actual outcome rates for each group. A model cannot simultaneously satisfy all three under most real-world data distributions. Choosing which metric to optimize requires an explicit decision about what fairness means in the deployment context, and that decision should be documented before the model is trained, not after it is evaluated. This survey of fairness concepts in machine learning provides the foundational taxonomy that the checklist items above build on.

Check 6: Regulatory Compliance Documentation

If the model falls under the EU AI Act’s definition of a high-risk AI system, which includes models used in employment, education, credit scoring, law enforcement, and several other categories, the compliance timeline is now settled: following the Digital Omnibus amendment formally adopted by the European Parliament and Council in June 2026, standalone Annex III high-risk AI systems must meet data governance and bias testing requirements by December 2, 2027. 

This is a deferral from the original August 2026 deadline, but the regulatory direction has not changed, and preparation is expected to be underway now. Article 10 of the EU AI Act specifies that training, validation, and testing datasets must be subject to data governance practices, must be relevant, representative, free of errors, and complete, with appropriate statistical properties for the specific population and context in which the system operates. Beyond fines, non-compliance creates a direct commercial risk: EU public procurement frameworks increasingly require AI Act compliance as a condition of tender eligibility, meaning a non-compliant system can disqualify an organization from public contracts before any fine is assessed.

What Remediation Actually Looks Like

Pre-Processing: Fix the Data Before Training

Pre-processing remediation addresses bias at the data level before training begins. The options include resampling underrepresented groups to bring their representation closer to the deployment distribution, reweighting training examples to increase the influence of underrepresented groups on model weights, and targeted data collection to fill coverage gaps identified in the representation audit. Pre-processing remediation is the most durable because it fixes the root cause rather than adjusting the model’s outputs downstream.

In-Processing: Constrain the Training

In-processing remediation adds fairness constraints to the training objective. This typically means adding a penalty term to the loss function that penalizes prediction disparity across demographic groups, or using an adversarial training approach where a separate model is trained to predict the demographic group from the primary model’s outputs. In-processing approaches require that demographic labels are available during training, which is not always the case.

Post-Processing: Adjust the Outputs

Post-processing remediation adjusts the model’s decision thresholds after training to equalize a chosen fairness metric across demographic groups. This is the easiest to implement and the most fragile, because it addresses the symptom rather than the cause. A threshold adjustment that achieves demographic parity on the evaluation set may not generalize to production traffic if the production distribution differs from the evaluation set. Post-processing remediation should be treated as a stopgap while pre-processing and in-processing remediation are implemented.

How Digital Divide Data Can Help

Digital Divide Data supports enterprise AI teams running data-level bias audits and implementing the remediation programs that audit findings require. For programs measuring representation gaps and label consistency across demographic subgroups, model evaluation services design evaluation frameworks disaggregated by the subgroups relevant to the deployment context rather than reporting only aggregate metrics. 

For programs that need targeted data collection to close coverage gaps identified in a representation audit, data collection and curation services source training examples from the underrepresented groups and contexts the audit identified. For programs addressing label-level bias through annotator calibration and guideline revision, trust and safety solutions provide annotation teams with calibration frameworks that measure and reduce subgroup-level annotation inconsistency.

If your model is in production and you haven’t run a data-level bias audit, you’re managing a risk you haven’t measured. Talk to an expert.

Conclusion

The six checklist items above are all data-level activities that need to happen before training and again after evaluation:

  • Representation audit
  • Label consistency audit
  • Curation audit 
  • Performance disparity measurement
  • Fairness metric selection
  • Regulatory compliance documentation

None of them require changes to the model architecture. All of them require discipline about what the training data actually contains and how it was produced.

The organizations that catch bias early are the ones that treat the audit as a standard step in the data program rather than a response to a production failure. What does your current training data pipeline document about the demographic distribution of the data that fed your last model?

References

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1-35. https://arxiv.org/abs/1908.09635

European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council (EU AI Act). Official Journal of the European Union. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT). https://arxiv.org/abs/2001.00973

Frequently Asked Questions

Q1. Is bias auditing the same as fairness testing?

They overlap but are not identical. Bias auditing is a broader process that identifies where bias entered the system, covering data collection, labeling, and curation. Fairness testing is a specific evaluation activity that measures whether the model’s outputs meet a chosen fairness criterion. You can run fairness testing without a bias audit, but the results will tell you that a problem exists without telling you where it came from or how to fix it. A full bias audit includes fairness testing as one component alongside the data-level checks that identify root causes.

Q2. Which fairness metric should we use?

There is no universally correct answer because different metrics operationalize different ethical concepts of fairness, and they can mathematically conflict with each other under real-world data distributions. The choice should be driven by the deployment context and the consequences of different error types for each affected group. A credit scoring model where false negatives disproportionately harm one group warrants a different metric than a content moderation model where false positives disproportionately silence one group. Document the choice and the reasoning before training begins, not after.

Q3. How often should a bias audit be run?

Before the first deployment of a model, whenever the training data is updated in a way that changes its composition, whenever the model is retrained or fine-tuned, and at a regular cadence after deployment, typically quarterly for high-stakes applications, to catch distribution drift in the production traffic that the original training set did not anticipate. One-time pre-deployment auditing is insufficient because deployment environments change and model behavior can drift as production traffic diverges from the training distribution.

Q4. What data is needed to run a demographic subgroup analysis?

Ideally, demographic attributes are captured at data collection and preserved through the annotation and curation pipeline so they are available for disaggregated analysis. When this is not the case, demographic attributes can be inferred using name-based classifiers, language model-based classifiers, or proxy variables that correlate with demographic characteristics. Inferred demographics introduce their own error rates and should be treated as approximate rather than definitive. For regulated applications where demographic analysis is required, the most defensible approach is to collect demographic attributes directly and with participant consent at the point of data collection.

Q5. Does a bias audit guarantee the model is fair?

No. A bias audit identifies measurable disparities in the training data and model outputs against specific metrics. It does not guarantee fairness in a philosophical or legal sense, because fairness is context-dependent and the audit’s conclusions are bounded by the metrics chosen, the subgroups analyzed, and the evaluation data used. What a thorough bias audit does provide is documented evidence of due diligence, specific findings that can be addressed through remediation, and a defensible record of what was measured and what was done about it. That is what regulators and enterprise governance programs require.

How to Audit an AI Model for Bias: A Practical Data-Level Checklist Read Post »

AI data pipeline services

The Enterprise Buyer’s Guide to AI Data Pipelines in 2026

AI data pipeline services are managed, end-to-end workflows that carry raw data through ingestion, transformation, labeling, validation, versioning, and delivery, so machine learning models receive training-ready inputs on a predictable schedule. For enterprise buyers in 2026, the real decision is whether to run this pipeline in-house or hand it to a managed provider that owns the human labeling and quality layer most teams underestimate. The right answer depends on data volume, domain complexity, regulatory exposure, and how much model accuracy rides on annotation quality.

Most AI programs stall in the same place. The architecture is sound, the compute is provisioned, and the pilot works on a curated sample; then production data arrives and the pipeline underneath cannot keep it clean, labeled, and versioned at volume. This is the gap that managed AI data pipeline services are built to close, and the strongest providers pair infrastructure with end-to-end data collection and curation. Buyers who understand what these services include and where they tend to fail are the ones who avoid paying for a pipeline that quietly produces unusable data.

Key Takeaways

  • AI data pipeline services are managed workflows that carry your raw data through collection, cleaning, labeling, checking, versioning, and delivery, so models always get data they can learn from.
  • The work runs in stages, and a weak stage quietly damages every stage after it, which is why quality has to be measured at each handoff rather than at the end.
  • Unlike a normal data pipeline that ends at a report a person reads, an AI pipeline feeds the model directly and loops back for retraining, so mistakes go straight into the model with no human to catch them.
  • Most AI projects fail because of the data underneath them, not the model on top, and the people who label and verify that data are the part companies most often underfund.
  • Building this in-house suits teams with rare, highly specialised data and deep existing expertise, while most enterprises get there faster with a managed or hybrid partner.
  • When comparing vendors, ask for proof including measured labeling accuracy, repeatable datasets, and clear security documentation, instead of trusting claims about quality.

What are AI data pipeline services?

AI data pipeline services are outsourced or co-managed programs that handle the movement, preparation, and quality control of the data feeding a machine learning system. They span the full path from source systems to model-ready datasets, and they usually bundle data engineering for AI with human annotation and validation. The term overlaps with related labels such as data operations and ML data preparation, but the scope stays consistent: get the right data, in the right shape, to the model, repeatedly and reliably. Reliable data pipelines are foundational elements for any AI system, and successful systems treat this as core infrastructure rather than a one-time project.

The distinction that matters for buyers is the one between a data pipeline (the technical plumbing) and AI data pipeline services (the plumbing plus the people and processes that keep the data trustworthy). A pipeline that moves data on schedule but delivers mislabeled or biased examples will train a model that fails in production. Gartner’s analysis of AI-ready data found that through 2026, organizations will abandon 60% of AI projects that lack properly prepared data, and that 63% of organizations either lack or are unsure of the data management practices AI requires. Those failures rarely trace back to the model itself.

This is why the field has shifted toward data-centric AI, where performance gains come from improving the data rather than re-architecting the model. A widely cited survey on data-centric AI describes training-data development, meaning collection, labeling, and preparation, as the primary lever for reliable model behavior. Managed pipeline services operationalize that idea. They wrap disciplined collection, annotation, and quality assurance around the data before it ever reaches training.

It helps to be concrete about what “AI-ready” means, because the phrase gets used loosely. Ready data is aligned to a specific use case, governed at the level of the individual data asset, produced by automated pipelines with quality gates, and quality-assured continuously rather than in periodic audits. Traditional data management runs on reporting cadences, where a quarterly review is fine. Models in production need quality signals measured in hours, and that mismatch is where most pipeline problems begin. A managed service exists to hold that continuous standard, so the internal team does not have to staff for it around the clock.

How does an AI data pipeline work, stage by stage?

An AI data pipeline is a sequence of stages, each with its own failure modes and quality gates. Weakness at any stage propagates downstream, so mature programs measure and control every handoff. The six core stages below describe what a well-run managed service actually delivers.

  1. Ingestion: Raw data is pulled from source systems such as sensors, logs, documents, databases, and third-party feeds, then normalized into a consistent format. Hybrid environments, where legacy on-premises systems sit beside cloud warehouses, are where ingestion most often breaks.
  2. Transformation: Data is cleaned, deduplicated, standardized, and enriched so downstream stages receive predictable inputs. Poor transformation lets duplicate or malformed records reach the model, which then learns patterns that do not exist.
  3. Labeling: Human annotators, often supported by pre-labeling models, add the ground-truth labels a supervised model learns from. This is the stage tooling-first vendors most often underinvest in, and multimodal data annotation across text, image, video, and sensor streams is where domain expertise earns its cost.
  4. Validation: Labeled data is checked for accuracy, consistency, and coverage before it is accepted. Inter-annotator agreement, gold-standard audits, and independent model evaluation turn “we labeled it” into “we can defend this label”.
  5. Versioning: Datasets, labels, code, and configurations are versioned so any training run can be reproduced and any regression can be traced to its source. Without versioning, a drop in model accuracy becomes an unsolvable mystery.
  6. Delivery: Model-ready datasets are handed to training and inference systems on a defined schedule, with quality and freshness service levels attached.

Between these stages sit data contracts, which are agreements about schema, freshness, and quality that each stage must meet before the next accepts its output. When a contract is violated, an alert fires before bad data reaches training. This is the difference between a pipeline that fails loudly and early and one that silently degrades a model over weeks. Strong managed services make these contracts explicit and measurable, so quality is a number on a dashboard rather than a matter of trust.

A production pipeline also includes a feedback loop. Model outputs are monitored, drift is detected, and fresh data is routed back through the same stages for retraining. The loop is what keeps a deployed model accurate as the real world changes around it. In sensor-heavy domains such as autonomous driving, that loop runs constantly because new edge cases appear in the field faster than any fixed dataset can anticipate.

What is the difference between a data pipeline and an AI pipeline?

A traditional data pipeline is a one-way street. It extracts data, transforms it, and loads it into a warehouse or dashboard, where a human reads the result. The pipeline’s job ends at delivery, and a person catches most errors before they cause harm.

An AI pipeline extends that path and closes it into a loop. It adds feature engineering, labeling, model training, and monitoring, then feeds model outcomes back to improve the next cycle. Because a model consumes the data directly, no human reads a dashboard to catch a bad batch, so quality control has to live inside the pipeline. Data orchestration for AI at scale becomes a first-class concern because dozens of stages, datasets, and model versions all have to stay coordinated.

An AI pipeline also introduces structures a reporting pipeline never needs, such as a feature store, which is a governed repository of the processed inputs a model consumes for both training and live inference. Keeping training features and serving features consistent is a problem business intelligence never had to solve, and getting it wrong produces models that score well in testing and fail in production. This is one more reason the AI pipeline demands tighter control than its reporting-era ancestor.

The other difference is standards. A dashboard tolerates a small share of dirty rows because a human discounts them at a glance. A model treats every example as truth and will happily learn from a mislabeled one. That raises the bar on labeling accuracy and validation far above what traditional business intelligence ever required.

How do you build a scalable AI data pipeline?

Scalability is decided early, in the design of the pipeline, and it cannot be bolted on once volume climbs. Teams that build for a pilot’s data volume usually rebuild within a year, because the tooling, quality process, and staffing that work for ten thousand examples collapse at ten million. Designing for the target volume from the start avoids that expensive second build.

Building a pipeline that holds up at scale rests on a few durable principles:

  • Standardize quality gates: Define accuracy thresholds, inter-annotator agreement targets, and freshness service levels, then enforce them automatically at each stage.
  • Version everything: Data, labels, code, and model configurations all need version control so results stay reproducible and regressions stay traceable.
  • Separate the human layer from the tooling layer: Annotation workforces and QA processes should scale independently of the ingestion and transformation stack.
  • Instrument for drift: Continuous monitoring of data and model behavior lets retraining trigger on evidence rather than on a fixed calendar.

The constraint most teams miss is trained people. A scalable pipeline needs a trained, managed annotation workforce with domain knowledge, and standing up that capability internally takes months. McKinsey’s 2025 State of AI survey found that 88% of organizations now use AI in at least one function, yet only about a third have scaled it enterprise-wide, and high performers are far more likely to have defined processes for when model outputs need human validation. The human quality layer, more than the algorithm, is what separates the two groups.

The cost of getting scalability wrong is technical debt that compounds. Data teams that spend most of their time maintaining fragile pipelines are firefighting rather than building, and every quarter of deferred quality work makes the eventual cleanup larger. Designing quality gates, versioning, and a managed workforce into the pipeline from day one is cheaper than retrofitting them once a model is already in production and already trusted by the business.

Managed service or in-house build: which fits your program?

The build-versus-buy decision turns on a few honest questions about cost, speed, and control. Building in-house makes sense when data is highly proprietary, the domain is narrow enough for a small expert team, and the organization already has data engineering and annotation management depth. For most enterprises, that combination is rare. The trade-offs of weighing a data annotation provider against an in-house team usually favor a managed or hybrid model once volume and domain breadth grow.

A managed service accelerates time-to-value and absorbs the operational burden of hiring, training, and retaining annotators. The common objections are real. Fully managed services can raise data-residency and control concerns in regulated industries, and some pricing models penalize scale. Those risks are manageable with the right contract terms, deployment model, and governance, which is why the vendor evaluation below matters as much as the build-versus-buy call itself.

A hybrid model is often the pragmatic answer. The enterprise keeps ownership of strategy, sensitive data, and final acceptance, while the provider runs collection, annotation, validation, and delivery at scale. This keeps control where it belongs and puts volume where it is cheapest to handle.

What should you look for in an AI data pipeline vendor?

Vendor selection is an architecture decision with long consequences, and a connector count on a slide tells you little about whether the data will be trustworthy. The questions that predict success are about the human quality layer, governance, and how a provider behaves when something breaks. The capabilities matrix below gives buyers a structured way to compare providers on what actually drives model performance.

Capability What strong looks like Warning sign
Data collection & curation Sourcing, cleaning, and curation run as a managed service with documented provenance Vendor only labels data you supply, with no curation
Annotation quality Measured inter-annotator agreement, gold-standard audits, domain-trained annotators “High quality” claimed with no metrics attached
Multimodal coverage Text, image, video, audio, and sensor data handled by one provider Single-modality shop staffing a multimodal program
Validation & evaluation Independent evaluation, plus bias and coverage checks before delivery QA limited to occasional spot checks
Versioning & reproducibility Datasets, labels, code, and configs versioned end-to-end No lineage; training runs cannot be reproduced
Governance & security RBAC, encryption in transit and at rest, audit trails, no training on your data Vague compliance badge with no documentation
Deployment model Cloud, hybrid, and on-prem options to fit data-residency rules Cloud-only in a regulated environment
Support & SLAs Documented response times, plus freshness and accuracy service levels SLAs “available on request,” never shown
Pricing predictability Transparent, volume-aware pricing Usage-based billing that punishes scale

For regulated industries, deployment models and compliance coverage often decide the shortlist before any other feature matters. A provider that is cloud-only cannot serve a program with strict data-residency rules, and a generic compliance badge is not the same as documentation you can hand to an auditor. Buyers in healthcare, finance, defense, and public sector should treat the deployment model and the governance posture as gating criteria, then compare on annotation quality and coverage within the providers that clear that bar.

The single most useful filter is evidence. A provider that can show measured annotation accuracy, reproducible datasets, and a documented governance posture is describing a program that will hold up in production. A fuller checklist for how to evaluate AI training data providers usually covers the diligence questions worth asking before a contract is signed.

How Digital Divide Data Can Help

Digital Divide Data runs the full AI data pipeline as a managed service, with the human quality layer built in rather than bolted on. Our teams handle end-to-end data collection and curation, multimodal annotation across text, image, video, audio, and sensor data, and the validation and versioning that keep datasets reproducible. This matters most in Physical AI, ADAS, and autonomous systems, where a single mislabeled sensor frame can propagate into a safety-relevant model error.

Where programs need an independent check on quality, our model evaluation services provide accuracy testing, bias and fairness assessment, and factual-consistency review before models reach production. We also support human preference optimization, red teaming, and trust and safety work, so the pipeline covers not only training data but the evaluation and alignment stages that decide whether a model behaves as intended. Our delivery model is designed to scale a trained, managed annotation workforce without forcing the enterprise to build that capability internally.

Build an AI data pipeline that delivers training-ready data you can actually trust. Talk to an Expert.

Conclusion

The organizations that get AI data pipeline services right in 2026 treat data quality as the core of the program, not a step to finish before the interesting work begins. They measure annotation accuracy, version their datasets, instrument for drift, and choose partners on evidence rather than connector counts. The organizations that get it wrong keep launching pilots on unprepared data and keep landing in the 60% of projects Gartner expects to be abandoned.

The pipeline underneath your model decides whether it scales or stalls. 

References

Gartner. (2025). Lack of AI-Ready Data Puts AI Projects at Risk. Gartner Newsroom. https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk

McKinsey & Company. (2025). The State of AI in 2025: Agents, Innovation, and Transformation. QuantumBlack, AI by McKinsey. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai

Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2025). Data-centric Artificial Intelligence: A Survey. ACM Computing Surveys. https://arxiv.org/abs/2303.10158

Frequently Asked Questions

What are AI data pipeline services in simple terms?

They are managed workflows that move your raw data through ingestion, transformation, labeling, validation, versioning, and delivery, so your machine learning models get clean, training-ready data on a reliable schedule. The provider usually handles both the technical plumbing and the human annotation and quality checks.

What is the difference between a data pipeline and an AI pipeline?

A regular data pipeline is a one-way street that ends at a dashboard a person reads, so a human catches most errors. An AI pipeline adds labeling, training, and monitoring, then loops model outcomes back for retraining, and because a model reads the data directly, quality control has to be built into the pipeline itself.

Should I build my AI data pipeline in-house or use a managed service?

Building in-house makes sense when your data is highly proprietary, your domain is narrow, and you already have data engineering and annotation management depth. For most enterprises, a managed or hybrid model is faster and cheaper once data volume and domain breadth grow, because standing up a trained annotation workforce internally takes months.

What should I look for in an AI data pipeline vendor?

Look for evidence rather than claims, including measured annotation accuracy, gold-standard audits, versioned and reproducible datasets, multimodal coverage, and a documented governance posture with encryption, access controls, and no training on your data. Also check that the deployment model and SLAs fit your industry’s data-residency and reliability requirements.

The Enterprise Buyer’s Guide to AI Data Pipelines in 2026 Read Post »

AI in Supply Chain

AI in Supply Chain: What Demand Forecasting and Logistics Models Need From Training Data

Kevin Sahotsky

Almost every supply chain leader I talk to is already running an AI pilot of some kind: demand forecasting, route optimization, inventory planning. Most of them are also quietly frustrated, because the pilot performed well in the demo and then underdelivered once it touched real operations. The model wasn’t wrong about the math. It was working from data that didn’t reflect the supply chain it was actually being asked to plan for.

This is particularly relevant for supply chain leaders, demand planning teams, and operations executives who are past the pilot stage and trying to figure out why their AI forecasting tool isn’t closing the gap they expected. The industry-wide numbers back this up. Most organizations plan to use AI for supply chain decisions within the next couple of years, but only a small fraction have a formal strategy for getting there, and the gap between adoption and actual readiness is almost always a data gap before it’s a model gap.

This blog covers what demand forecasting and logistics models actually need from their training data to perform reliably in production, not just in a pilot. Data collection and curation services and AI data preparation services are the two capabilities most directly involved in closing the gap between a forecasting model that looks good on a slide and one that actually holds up against real demand volatility.

Key Takeaways

  • Demand forecasting models trained only on historical sales data systematically underperform during demand shifts, because the signal that predicts a shift rarely lives in the sales history itself.
  • Supply chain AI needs data integrated across systems that were never designed to talk to each other. Partner data chaos, not model architecture, is the most common reason forecasting and logistics AI underdelivers.
  • SKU-level and category-level forecasting have very different data requirements, and treating them the same way is one of the most common planning mistakes.
  • Exception and disruption data- the supplier delay, the port closure, the demand spike- is the training signal that determines whether a model can do more than predict business as usual.
  • Human review at the exception layer is what keeps automated forecasting accurate, because full autonomy isn’t the goal right now. Appropriate autonomy is.

Why Forecasting Models Underdeliver Outside the Pilot

Historical Sales Data Is a Starting Point, Not a Foundation

Traditional forecasting leaned almost entirely on historical sales data, and that’s exactly where a lot of AI forecasting pilots still start. The problem is that historical sales data tells you what happened under the conditions that existed at the time. It doesn’t tell you why those conditions are about to change. A model trained purely on sales history will perform reasonably well during stable periods and fail exactly when you need it most, during a demand shift, a new product launch, or a market disruption.

This isn’t a hypothetical concern. Industry data shows AI-powered forecasting can reduce forecast errors meaningfully and cut inventory costs, but those gains depend on the model having access to a broader mix of signals than historical sales curves alone. Retailers that combined external signals with real-time inventory visibility saw the greatest improvements, specifically because the model had something other than the past to reason from.

The Real Bottleneck Is Partner Data Chaos

Ask supply chain leaders what’s actually holding AI back day to day, and the answer that comes up again and again isn’t the model. It’s the mess of formats, systems, and partner data that the model has to be fed from. Suppliers report inventory differently. Carriers report transit status on different schedules. Internal systems were built for different purposes at different times and were never designed to be queried together. Data engineering for AI that builds the integration layer connecting these disparate sources into a consistent, queryable structure is what turns partner data chaos into something a forecasting model can actually use, and it is consistently the unglamorous work that determines whether the visible AI layer performs.

What Demand Forecasting Models Actually Need

SKU-Level vs. Category-Level Forecasting Have Different Data Needs

One of the most common mistakes I see is treating SKU-level and category-level forecasting as the same data problem at different resolutions. They aren’t. Category-level forecasting can tolerate more noise in any individual data point because the aggregation smooths it out. SKU-level forecasting, especially for products with intermittent or erratic demand patterns, needs cleaner, more granular data because there’s no aggregation to hide a labeling error or a missing data point.

This matters most for businesses managing SKU proliferation: large retailers and consumer goods companies that are tracking demand across thousands of individual products. A forecasting approach that works fine at the category level can produce confidently wrong SKU-level forecasts if the underlying data wasn’t curated with that level of granularity in mind from the start.

External Signals Are Not Optional Anymore

The forecasting approaches that are actually moving the needle right now combine internal sales data with external signals: economic indicators, weather patterns, regional events, competitor activity, and social signals where relevant. Collecting and structuring these external signals consistently, so they can be joined to internal sales data on a common timeline, is a data engineering task that most internal teams underestimate the effort of. Data collection and curation services that source and standardize external demand signals on an ongoing basis, not as a one-time enrichment, are what let a forecasting model actually use this information rather than treating it as an occasional input that goes stale.

Seasonality and Intermittent Demand Need Explicit Handling

Demand patterns that are seasonal, intermittent, or erratic break the assumptions that simpler forecasting methods rely on. A model that hasn’t been given enough historical cycles to learn a seasonal pattern, or training data with sparse and irregular intermittent-demand examples, will produce point forecasts that look plausible and are systematically wrong in predictable ways: missing the seasonal peak, or smoothing over the spikes that intermittent-demand products actually exhibit. The fix isn’t a different algorithm. It’s making sure the training data includes enough cycles and enough representation of the demand pattern types the business actually has.

What Logistics and Routing Models Need

Real-Time Data, Not Just Planning Data

Route optimization and ETA prediction depend on data that’s current, not just historical. A model trained on historical transit times without real-time traffic, weather, and carrier status data will optimize for a world that no longer exists by the time the truck leaves the dock. The practical implication is that logistics AI needs a live data pipeline, not a periodically refreshed training set, and the infrastructure to keep that pipeline current is a meaningfully different investment than the one-time data preparation that a static forecasting model might get away with.

Exception Data Is the Most Valuable and Least Collected

Most logistics data pipelines are built to capture the normal case well and the exception case poorly. The supplier delay, the port closure, the carrier capacity shortfall- these are exactly the events that determine whether a logistics AI system adds value beyond what a simple rules engine could already do, and they’re also the events most likely to be missing, inconsistently labeled, or buried in free-text notes rather than structured fields. AI data preparation services that specifically target exception event extraction and structuring, pulling disruption data out of free text and into a consistent schema, give logistics models the training signal they need to do more than optimize for business as usual.

Why Human Review at the Exception Layer Still Matters

Full autonomy in supply chain AI isn’t where the industry actually is right now, and the practitioners closest to deployment are honest about that. The current consensus across the field is that appropriate autonomy, not full autonomy, is the right target for 2026. Automated forecasts paired with human review on exceptions and material categories consistently outperform either fully automated or fully manual approaches.

Building that human review layer into the data pipeline, not as an afterthought but as a designed checkpoint, is what keeps a forecasting system’s error rate from compounding silently. Model evaluation services that score forecast accuracy by category, by exception type, and by demand pattern, rather than as a single aggregate accuracy number, are what let a supply chain team know where the human review needs to be concentrated rather than spread thin across everything.

How Digital Divide Data Can Help

Digital Divide Data supports supply chain and logistics teams building the data foundation that demand forecasting and routing models actually need. For programs that need external demand signals collected and standardized on an ongoing basis, data collection and curation services source and structure economic, weather, and market signals so they can be joined cleanly to internal sales data. 

For programs that need exception and disruption events extracted from free-text logs into structured, model-ready fields, AI data preparation services turn unstructured supplier, carrier, and operations notes into the training signal that logistics models need to handle disruption. For programs connecting fragmented partner and internal systems into a single queryable pipeline, data engineering for AI builds the integration layer that turns partner data chaos into a usable forecasting input.

If your forecasting model performs well in the pilot and underdelivers in production, the gap is almost always in the data feeding it, not the model architecture. Talk to an expert.

Conclusion

The supply chain AI gap that emerges between a strong pilot and a disappointing production rollout is rarely an algorithmic problem. It’s a data problem: historical sales data without external signals, fragmented partner systems never designed to be queried together, and exception events that occur in the operation but never make it into a structured training set. Each of these is solvable, but only if the team treats data integration and curation as the primary investment rather than something the model is supposed to work around.

The organizations pulling ahead in supply chain AI aren’t the ones with the most sophisticated forecasting algorithm. They’re the ones that did the less visible work of making sure their models had real, current, well-structured signal to learn from. What does your current forecasting pipeline actually feed the model, and how much of it is historical sales data alone?

References

Logistics Viewpoints. (2025, December 22). AI in logistics: What actually worked in 2025 and what will scale in 2026. https://logisticsviewpoints.com/2025/12/22/ai-in-logistics-what-actually-worked-in-2025-and-what-will-scale-in-2026/

Inbound Logistics. (2026, January 8). AI in supply chain management: 2026 outlook. https://www.inboundlogistics.com/articles/ai-in-supply-chain-management-how-useful-will-it-be-in-2026/

Frequently Asked Questions

Q1. Why does a demand forecasting model that performed well in a pilot underdeliver once it is deployed at scale?

Pilots are often run on a clean, curated slice of data and a stable demand period. Production exposes the model to the messier reality: fragmented partner data, demand patterns the pilot dataset didn’t include, and exception events that weren’t part of the pilot’s scope. The model’s architecture usually isn’t the problem. The training data it’s actually getting in production is narrower or noisier than what it learned from during the pilot, and that gap is what shows up as underperformance.

Q2. What external data signals matter most for demand forecasting beyond historical sales?

It depends on the category, but the signals that consistently add value are economic indicators relevant to the customer base, weather data for weather-sensitive categories, regional event calendars, and competitor pricing or promotion activity where it’s trackable. The specific mix matters less than having a consistent process for collecting and standardizing whichever signals are relevant to your categories, so the model can actually learn a stable relationship between the signal and the demand shift rather than seeing it inconsistently.

Q3. How should a supply chain team prioritize data investment between forecasting accuracy and logistics optimization?

Start with whichever side is generating the more expensive errors right now. If you’re consistently overstocking or understocking specific categories, the forecasting data investment will pay off faster. If you’re missing delivery windows or absorbing avoidable transportation costs because of routing decisions made on stale data, the logistics data pipeline is the higher-value investment. Most teams need both eventually, but sequencing the investment around your most expensive current error avoids spreading a limited budget too thin to fix either one well.

Q4. How much human review should remain in an automated forecasting and logistics pipeline?

Enough that exceptions and high-consequence categories get a human check before the system acts on them automatically. Full autonomy isn’t where the field is right now, and the practitioners closest to production deployment are explicit that appropriate autonomy, not full autonomy, is this year’s realistic target. A practical approach is to automate the routine, high-confidence cases and route anything flagged as an exception, a material category, or a low-confidence prediction to a human reviewer before it triggers a downstream action.

Q5. What is the most common reason a supply chain AI program stalls after the pilot phase?

Underestimating the data integration work required to move from a pilot dataset to a production data pipeline. A pilot can run on a manually assembled, cleaned dataset. Production requires an ongoing pipeline that ingests, standardizes, and validates data from multiple internal systems and external partners on a continuous basis. Teams that scope the pilot but not the production data infrastructure consistently find that the second phase takes longer and costs more than the first, and that gap is where many programs stall.

AI in Supply Chain: What Demand Forecasting and Logistics Models Need From Training Data Read Post »

Evaluate VLA Model

How to Evaluate VLA Model for Real-World Deployment: Grounding, Planning, and Action Fidelity

Kevin Sahotsky

Here’s a question I get from robotics and physical AI teams more often than I used to: we have a VLA model that looks impressive in the demo, how do we know if it will actually hold up once it leaves the lab? It’s a fair question, and the honest answer is that most teams do not have a good way to answer it yet. The benchmarks the model was trained and reported against often measure something narrower than what deployment actually requires.

Vision-Language-Action models are being evaluated in the way earlier generations of computer vision models were evaluated: a held-out test set, a success rate, and a leaderboard position. That approach tells you how the model performs on a distribution similar to its training data. It tells you very little about whether the model will ground a spoken instruction correctly in a cluttered warehouse, plan a multi-step task when the first attempt fails, or execute an action with the physical fidelity that a real task requires. This is particularly relevant for robotics program leads, physical AI product teams, and operations leaders evaluating whether a VLA model is ready to move from a controlled pilot into a live deployment.

This blog walks through the three capabilities that actually determine whether a VLA model is deployment-ready: grounding, planning, and action fidelity, and what evaluating each of them looks like in practice. Model evaluation services and video annotation services are the two capabilities most directly involved in building VLA evaluation programs that predict real-world performance rather than benchmark performance.

Key Takeaways

  • Standard VLA benchmarks measure performance on a distribution similar to the training data. They do not reliably predict performance in a specific deployment environment with its own object sets, lighting, and task variations.
  • Grounding, planning, and action fidelity are three distinct capabilities that fail independently. A model can ground language well and still fail at multi-step planning, or plan well and still execute with poor physical fidelity.
  • Out-of-distribution evaluation, testing on object placements, lighting, and task variations the model has not seen, is a better predictor of deployment performance than in-distribution benchmark scores.
  • Action fidelity cannot be assessed from success rate alone. Two policies with the same success rate can have very different margins for error, and that margin is what determines reliability at scale.
  • A model evaluation program built around your specific deployment task taxonomy will catch failure modes that a general VLA leaderboard never surfaces.

Why Standard Benchmarks Undersell the Real Question

What Leaderboard Scores Actually Measure

Most published VLA benchmarks evaluate models on tasks and environments that are either simulated or closely matched to the training distribution. A model can score well on these benchmarks because the test conditions are similar enough to what it has already seen. That is a legitimate measure of in-distribution capability. It is not a measure of whether the model will work in your warehouse, your kitchen, or your assembly line, where the objects, the lighting, and the task variations will not match the benchmark distribution.

This gap matters more for physical AI than it did for earlier generations of language or vision models, because the cost of a wrong answer is physical. A chatbot that gives an unhelpful response is an inconvenience. A robot that misjudges a grasp or executes the wrong action in a live environment is a safety and operational problem.

Out-of-Distribution Performance Is the Real Signal

Recent benchmarking work in the field has made a point that matters here: models trained primarily on action data can transfer reasonably well to environments that resemble their training distribution, but performance drops sharply when visual conditions or task mechanics shift outside that distribution. That is the gap that matters for deployment. If your environment, your object set, or your task structure differs meaningfully from what the model was trained on, the benchmark score tells you very little about what will happen in production.

The practical implication is that evaluation needs to be built around your specific deployment context, not borrowed wholesale from a published leaderboard. A model that ranks well on a general benchmark may still fail consistently on the specific variations your environment introduces.

Grounding: Does the Model Understand What You Are Asking?

What Grounding Failures Look Like

Grounding is the model’s ability to connect a language instruction to the correct object, location, or action in its visual field. A grounding failure looks like the model picking up the wrong object when two similar items are present, or misinterpreting a spatial reference like “the one on the left” when the scene has shifted from how it appeared in training.

Grounding failures are often invisible in simple test environments because there is only one plausible object or location for the model to act on. They become visible the moment you add visual clutter, similar-looking objects, or ambiguous spatial language, which is exactly what real environments contain in abundance.

The business cost of a grounding failure shows up as rework and damaged trust rather than a single dramatic incident. A model that picks up the wrong part on an assembly line creates a defect that gets caught downstream, at a higher cost than catching it at the source. A model that misreads a spatial instruction in a fulfillment center sends the wrong item, which becomes a customer-facing SLA breach and a return to process. Multiply a small grounding error rate by daily production volume, and the cost stops looking small.

Evaluating Grounding Under Realistic Ambiguity

A grounding evaluation needs deliberately ambiguous scenes: multiple objects of similar type, instructions that require spatial or relational reasoning, and language phrasings that vary from the canonical form the model may have been trained on. Video annotation services that label ground-truth object references and spatial relationships in evaluation footage give you the basis for scoring whether the model’s grounding matches what a human would understand the instruction to mean, rather than just whether the model picked up some object.

Planning: Does the Model Handle Multi-Step Tasks and Recover From Failure?

Single-Step Success Hides Planning Weakness

Many real tasks require a sequence of actions, and a model can execute each action competently while still failing at the task because it does not sequence them correctly, does not recognize when an earlier step failed, or cannot adapt the plan when the environment does not match its expectation. A model evaluated only on isolated single-step actions will look far more capable than it will behave in a multi-step task.

Hierarchical approaches that separate high-level planning from low-level execution have shown that grounding ambiguous instructions and adapting plans dynamically remains one of the harder open problems in the field. That should inform how much weight you put on a model’s single-step benchmark score relative to its actual planning behavior.

Planning failures are typically what cause unplanned downtime, not a single bad action. A robot that cannot recognize a failed step will either stall and wait for human intervention, which stops the line, or continue executing a plan built on a false assumption, which can damage product or equipment before anyone notices. Both outcomes carry a direct cost in lost throughput, and the second carries an added repair or scrap cost on top of it.

Designing an Evaluation for Recovery Behavior

Planning evaluation should deliberately introduce failure points: an object that is not where the model expects it, a step that cannot be completed on the first attempt, or a change in the environment mid-task. The question is not whether the model can execute a clean multi-step task when everything goes as expected. It is whether the model notices when something has gone wrong and adapts rather than continuing to execute a plan based on a stale assumption.

Action Fidelity: How Precisely Does the Model Execute?

Why Success Rate Alone Is Not Enough

Two policies can report the same success rate on a benchmark while having very different margins for error. One policy might complete a grasp with a wide, stable margin every time. Another might complete the same grasp at the edge of what is mechanically possible, succeeding in the test conditions but failing the moment an object’s weight, texture, or position shifts slightly. Success rate does not distinguish between these two cases, and that distinction is exactly what determines whether a model is reliable at production scale.

Action fidelity evaluation requires looking past the binary success label to the quality of the execution itself: trajectory smoothness, contact stability, and how close the action came to the failure boundary, even when it technically succeeded.

A narrow action fidelity margin is the kind of risk that does not show up until volume increases or conditions drift slightly, and then it shows up as a safety incident or an equipment damage claim rather than a quality metric. A grasp that succeeds at the edge of mechanical stability in a pilot of fifty units can fail consistently at a production volume of five thousand, once object weight or surface friction varies even slightly from the pilot batch. That is the gap between a model that looked ready in evaluation and one that was not.

Building Action Fidelity Into the Evaluation Protocol

This requires frame-level review of execution quality, not just episode-level success labels. Model evaluation services that score action fidelity on dimensions like grasp stability margin and trajectory precision, not just task completion, surface the difference between a model that succeeds reliably and one that succeeds narrowly.

Building an Evaluation Program Around Your Deployment, Not the Leaderboard

The most useful thing you can do before deploying a VLA model is to define your own task taxonomy: the specific objects, environments, instruction phrasings, and failure scenarios your deployment will actually involve. Then evaluate the model against that taxonomy directly, rather than relying on how it ranks on a general benchmark.

This is not a one-time gate before launch. Models get updated, deployment environments evolve, and new task variations show up that your original evaluation set did not anticipate. Data collection and curation services that continuously sample new deployment scenarios into your evaluation set keep the evaluation program honest as the deployment context changes.

Signs Your Current Evaluation Is Not Enough, and Whether to Build or Buy

Not every team is at the same starting point, and it is worth being honest about where you actually are before investing further. A few signs your current evaluation program is not enough: your only performance number comes from a published benchmark or the model provider’s own reported metrics; you have never tested the model against object types, lighting, or instruction phrasings specific to your facility; your evaluation set has not changed since the pilot, even though your deployment environment has; or you are relying on field incident reports, rather than a structured evaluation process, to tell you when something is wrong.

If two or more of these are true, the question becomes whether to build this evaluation capability in-house or bring in a partner to run it. Building in-house makes sense if you already have ML engineers who understand evaluation design, your deployment environment is stable enough that a one-time investment in tooling will keep paying off, and you have the headcount to maintain the evaluation set as conditions change. Buying makes sense if your team’s strength is in the application and the robotics integration rather than in evaluation methodology, if your deployment environment is still evolving and the evaluation set will need frequent updates, or if you need this running before your next deployment milestone and do not have the lead time to build the capability from scratch. Most teams that choose to buy are not outsourcing judgment; they are outsourcing the ongoing labor of keeping an evaluation set current, which is the part that erodes fastest when left to a part-time internal owner.

How Digital Divide Data Can Help

Digital Divide Data supports robotics and physical AI teams building VLA evaluation programs that are grounded in the specific environments and tasks those models will face. For programs designing grounding and action fidelity evaluations, model evaluation services build evaluation frameworks around your deployment task taxonomy, with scoring dimensions that go beyond binary success rate to capture execution quality and failure margins. 

For programs that need labeled ground truth for grounding and planning evaluation, video annotation services provide annotation of object references, spatial relationships, and task phase structure in evaluation footage. For programs that need to keep their evaluation sets current as deployment environments evolve, data collection and curation services continuously source new evaluation scenarios from the field rather than relying on a static benchmark.

If your VLA evaluation program is built around a published benchmark rather than your actual deployment task taxonomy, you will not see the failure modes that matter until they show up in production. Talk to an expert.

Conclusion

A VLA model that looks strong on a published benchmark can still fail in your specific deployment, because the benchmark was never designed to predict performance in your environment. Grounding, planning, and action fidelity are three distinct capabilities that each fail in their own way, and a benchmark score that averages across all three will hide exactly the failure you need to catch before deployment.

The teams that get this right build their own evaluation taxonomy around the objects, environments, and task variations their deployment will actually involve, and they keep updating it as conditions change. What does your current VLA evaluation actually tell you about how the model will behave in your specific environment, not a benchmark’s?

References

Guruprasad, P., Wang, Y., et al. (2025). Benchmarking the generality of vision-language-action models. https://arxiv.org/abs/2512.11315

Li, X., Hsu, K., Gu, J., Pertsch, K., Mees, O., Walke, H. R., Fu, C., Lunawat, I., Sieh, I., Kirmani, S., et al. (2024). Evaluating real-world robot manipulation policies in simulation. arXiv. https://arxiv.org/abs/2405.05941

Zhou, J., Ye, K., Liu, J., Ma, T., Wang, Z., Qiu, R., Lin, K., Zhao, Z., & Liang, J. (2025). Exploring the limits of vision-language-action manipulations in cross-task generalization. arXiv. https://arxiv.org/abs/2505.15660

Frequently Asked Questions

Q1. How is evaluating a VLA model different from evaluating a standard computer vision model?

A vision model is typically evaluated on a single capability, like classification or detection accuracy, against a static test set. A VLA model has to be evaluated across three interacting capabilities at once: whether it understands the instruction, whether it plans the right sequence of actions, and whether it executes those actions with enough physical precision to succeed. A model can be strong in one of these and weak in another, and a single aggregate success rate will not tell you which one is the problem.

Q2. What does an out-of-distribution evaluation set actually look like for a VLA model?

It is a test set built deliberately to differ from the model’s likely training distribution: object types it has not seen paired with familiar ones, lighting and background conditions different from the training environment, and instruction phrasings that vary from the canonical form. The goal is not to make the test unfairly hard. It is to find the boundary of where the model’s competence actually stops, which a test set drawn from the same distribution as the training will not reveal.

Q3. How do you evaluate a model’s ability to recover from a failed step in a multi-step task?

Build evaluation scenarios that deliberately introduce a failure partway through a task: move an object slightly, interrupt the action, or change the environment mid-sequence. Then assess whether the model recognizes that the expected state did not occur and adapts, or whether it continues executing a plan based on its original assumption. This requires reviewing the full episode, not just the outcome, because a model can recover successfully through an inefficient path or fail silently while still producing a result that looks plausible at a glance.

Q4. What is a reasonable success rate to expect from a VLA model before deployment?

There is no single universal threshold, because the right number depends on the cost of failure in your specific task, but rough industry ranges give you a starting anchor. Low-stakes sorting or bin-picking tasks with cheap recovery from a miss are often deployed in the 90 to 95 percent success range, with a human or a simple fallback catching the rest. Tasks involving variable or fragile objects, such as warehouse pick-and-pack with mixed SKUs, generally need to clear 95 to 98 percent before the rework cost stops eating the labor savings. Tasks operating near people, expensive equipment, or in safety-relevant contexts, such as collaborative assembly or surgical-adjacent applications, are typically held to 99 percent or higher, often paired with a hard mechanical or software safety layer rather than relying on the model’s success rate alone. These are starting anchors, not certifications. What matters more than clearing a number is understanding the failure modes behind whatever rate you observe: whether failures are concentrated in specific object types, specific instruction phrasings, or specific task phases. That breakdown tells you whether the gap is fixable with more targeted data or whether it reflects a more fundamental limitation.

How to Evaluate VLA Model for Real-World Deployment: Grounding, Planning, and Action Fidelity Read Post »

Data Annotation Provider Pricing Models Decoded: Per-Label, Per-Hour, or Outcome-Based?

Data Annotation Provider Pricing Models Decoded: Per-Label, Per-Hour, or Outcome-Based?

Most data annotation providers price work in one of three ways: per-label (a fixed rate per annotation unit), per-hour (time-and-materials for annotator time), or outcome-based (payment tied to a quality SLA such as accuracy or acceptance rate). Per-label rewards volume and fits high-volume, well-specified tasks; per-hour fits complex or evolving work where time per item is hard to predict; outcome-based aligns the provider with the quality your model actually needs. The cheapest headline rate is rarely the cheapest total cost, because rework, rejected batches, and re-labeling are billed somewhere downstream.

A quote that looks inexpensive per label can become the most expensive option once those hidden costs are counted. Pricing structure is not just a budget line; it sets the incentives that shape throughput, quality, and how much oversight your own team has to supply. That makes the structure behind a quote worth as much scrutiny as the number itself, especially for programs that depend on large-scale data collection and curation and on multimodal annotation spanning image, video, text, audio, and sensor data.

Key Takeaways

  • Data annotation providers usually charge in one of three ways; a set fee per label, an hourly rate for time worked, or a price tied to the quality they deliver.
  • Per-label pricing is easy to budget and works best for large, simple, repetitive jobs.
  • Hourly pricing fits complex or changing work where you can’t predict how long each item will take.
  • Outcome-based pricing costs more upfront but is the only model that pays for results instead of just effort.
  • The lowest quoted price is often the most expensive once you add in fixing and redoing poor-quality labels.
  • To compare vendors fairly, look at the cost per usable label, not the headline rate.

How do data annotation companies charge?

A data annotation provider or an annotation partner turns raw data into labeled training data for AI and ML systems. Its charges usually bundle some mix of annotation, quality assurance, tooling, project management, and, in mature engagements, downstream model evaluation that confirms the labels actually improve model behavior. Three pricing structures dominate the market: per-label (per-unit), per-hour (time-and-materials), and outcome- or SLA-based contracts. Each one moves risk between buyer and vendor in a different direction.

Published market guides put simple bounding-box labeling at roughly three to eight cents per object, with dense segmentation and 3D work costing far more, while hourly rates range from about four to sixty dollars depending on region and annotator skill. Those ranges are useful for sanity-checking a quote, but they do not tell you which structure protects you. The deeper question is which model ties what you pay to what you can actually use in training. Researchers studying high-stakes AI documented data cascades, downstream failures triggered by upstream data problems, and affecting 92% of the practitioners they surveyed, which is why pricing that quietly trades away quality rarely saves money.

What is per-label pricing, and when does it work?

Per-label pricing charges a fixed amount for each annotation unit: a bounding box, a polygon, a labeled entity, or a transcribed segment. It is the easiest model to forecast, because cost scales directly with volume, and it fits high-volume, well-specified, repetitive work where the time per item barely varies. Teams budgeting bounding box annotation cost for large image sets usually find that per-label is the most transparent option.

The structure rewards throughput, which is exactly its risk; when annotators are paid per unit, speed can quietly win over care, and ambiguous edge cases get the fastest defensible label rather than the correct one. Per-label also says nothing about quality on its own; unless your contract specifies payment on accepted units, you can pay for labels that later fail review. The economics shift further as hybrid human-and-AI labeling moves routine units to model pre-labeling and reserves people for correction, which lowers per-unit cost but concentrates the hard, judgment-heavy cases that per-label rates often underprice.

When is per-hour (time-and-materials) pricing the right model?

Per-hour pricing, also called time-and-materials, bills for the actual time annotators and reviewers spend. It suits complex or evolving works; detailed semantic segmentation, multi-step NLP, sensor-fusion labeling, etc., where time per item swings too much for a per-unit rate to be fair to either side. It is also the sensible choice early in a project, when the specification is still moving, and you do not yet know how long each item takes.

The trade-off is predictability; total cost is hard to forecast, and the model can reward slowness unless you track output. The protection is through, agree on expected units per hour from comparable past work, and review it regularly. Teams looking to speed up a data annotation project usually find the lever is workflow design and tooling, not simply paying for more hours, so per-hour contracts work best when paired with reported productivity targets.

How does outcome-based and SLA pricing change vendor incentives?

Outcome-based pricing ties payment to measurable results, usually an accuracy threshold, an accuracy threshold, an inter-annotator agreement target, an acceptance rate, or a turnaround SLA. It usually arrives as a managed service or project fee that bundles annotation, QA, tooling, and reporting. On a per-unit basis, it is often the most expensive structure, but it is the only one that aligns the provider’s incentive with the quality your model actually needs.

That alignment matters because label quality is not cosmetic. A study of ten widely used benchmark datasets found pervasive label errors averaging around 3.4%, and correcting them was enough to change which model ranked best. If a few percent of label noise can reorder model rankings, a pricing model that pays only for accepted, audited output is buying something a per-label discount cannot. Understanding what a figure like 99.5% accuracy means in production is what makes an outcome SLA enforceable rather than decorative.

What are the red flags in a data annotation pricing proposal?

A proposal’s structure reveals more than its rate. A few patterns consistently signal trouble, and most trace back to unreliable annotation that has to be fixed later:

  1. A headline rate with no acceptance definition: If the quote does not say what counts as an accepted label, rework is being priced as your problem, not the vendor’s.
  2. Quality assurance billed as an extra: QA, gold-standard audits, and review passes are part of producing usable data, not an upsell. Watch for setup, onboarding, and rework fees that surface only after signing.
  3. No throughput or quality baseline: Per-hour quotes without expected units per hour, and per-label quotes without a stated quality bar, leave you unable to predict either cost or outcome.
  4. Rush and change-order premiums left vague: Expedited work legitimately costs more, but undefined premiums of thirty to fifty percent can swamp the base rate.
  5. Lock-in disguised as low pricing: A cheap rate on a proprietary platform you cannot export from raises your switching cost later.

How to compare data annotation vendor quotes fairly?

The mistake in most comparisons is treating per-label, per-hour, and managed-service quotes as if they measure the same thing and deliver the same accuracy & quality, but they don’t. The only fair basis is cost per accepted, usable unit, i.e., the total fee divided by the labels that survive quality bar, rework included. A disciplined approach to evaluating AI training data providers should normalize every quote to that number before any rate is compared.

To compare quotes on equal footing:

  • Convert each quote to a blended cost per accepted unit using a shared sample task and identical acceptance criteria.
  • Ask for throughput and quality figures from comparable completed projects, not stated rates alone.
  • List every add-on: tooling, QA, project management, compliance, expedited delivery, and fold it into the unit cost.
  • Run a paid pilot on the same gold-standard set so each vendor is measured against the same ground truth.

How Digital Divide Data Can Help

Digital Divide Data structures pricing around the unit that matters to your model; accepted, audited output. Its data collection and curation services build quality gates, inter-annotator agreement tracking, and gold-standard auditing into the workflow rather than billing them as afterthoughts, so the number you compare is the number you can train on.

For teams weighing outcome-based contracts, DDD’s model evaluation services connect annotation quality to measured downstream behavior, which is what makes an accuracy or acceptance SLA enforceable. Engagements scale from per-unit work on well-specified tasks to managed delivery on complex multimodal and sensor data, with the pricing structure chosen to fit the program rather than the other way around.

Compare annotation quotes on what your model can actually use, not the headline rate. 

Conclusion

Pricing structure is a decision about incentives, not just budget. Per-label rewards volume, per-hour rewards time, and outcome-based rewards, the quality your model depends on; the right choice follows from how stable your specification is and whether you can measure quality at acceptance. Organizations that normalize every quote to cost per accepted unit consistently spend less over a program’s life than those anchored to the lowest headline rate, because they stop paying twice for the same labels.

The teams that get this right treat the contract as part of their quality system; the ones that do not tend to discover the true cost during a mid-project scramble. 

References

Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. (2021). “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3411764.3445518

Northcutt, C. G., Athalye, A., & Mueller, J. (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. arXiv preprint arXiv:2103.14749. https://arxiv.org/abs/2103.14749

Frequently Asked Questions

How do data annotation companies charge?

Most use one of three structures: per-label (a set fee per annotation unit), per-hour (time-and-materials for annotator time), or outcome-based contracts tied to a quality SLA. The charge usually also folds in QA, tooling, and project management.

Is hourly or per-label annotation pricing better?

Neither is better in the abstract. Per-label fits high-volume, well-specified tasks where time per item barely varies, while per-hour fits complex or evolving work where you cannot predict how long each item takes. The deciding factor is how stable your specification is.

What should be in a data annotation pricing proposal?

A clear definition of what counts as an accepted label, QA and audit steps included rather than billed as extras, a throughput or quality baseline, and defined rush and change-order premiums. Missing any of these usually means rework is being priced as your problem.

How do I compare data annotation vendor quotes fairly?

Convert every quote to cost per accepted, usable unit using the same sample task and acceptance criteria, fold in all add-on fees, and run a paid pilot against an identical ground truth. Comparing headline rates alone hides where the real cost sits.

Data Annotation Provider Pricing Models Decoded: Per-Label, Per-Hour, or Outcome-Based? Read Post »

AI Evaluation Program

Why Your AI Evaluation Program Is Missing Cultural Failures, and How to Fix It

Kevin Sahotsky

Here’s a pattern I’ve seen more than once. An enterprise buys access to a frontier model, runs it through internal evaluations, and the results look good. Strong accuracy. Coherent outputs. The team gets comfortable. Then the model enters a customer-facing workflow serving users in the Middle East, Southeast Asia, or Sub-Saharan Africa, and something goes wrong. The outputs are technically correct in a narrow sense but contextually off. Users notice.  This is particularly relevant for AI procurement leads, product teams, and enterprise buyers deploying models in global or multilingual markets.

The evaluation wasn’t wrong. It was just evaluating the wrong thing. Standard benchmarks are predominantly designed around Western, English-language contexts. They measure capability on the kinds of inputs those contexts generate. When the deployment context is different, the benchmark stops being a reliable predictor of real-world performance.

Cultural alignment is becoming a first-order evaluation problem for any enterprise deploying AI in global markets. Model evaluation services and low-resource language services are the two capabilities most directly involved in closing the gap between what standard benchmarks measure and what global deployment actually requires.

Key Takeaways

  • Frontier models are trained predominantly on Western, English-language data. This produces systematic gaps in cultural knowledge, values alignment, and contextual reasoning that standard benchmarks do not surface.
  • Cultural failure is not a language problem. A model can be fluent in Arabic or Hindi while still applying Western cultural assumptions to content produced in those languages.
  • Standard benchmarks do not catch cultural misalignment. Evaluation programs that rely on existing leaderboard benchmarks will miss the failure modes that matter most in global deployments.
  • The evaluation gap is measurable. Culturally grounded human evaluation of production-representative inputs is the only reliable way to understand how a model will perform in a specific cultural context before that context reveals the failure.
  • The fix requires both better evaluation data and better training data. Identifying cultural gaps through evaluation and then closing them through targeted data collection are two sides of the same coin.

Why Frontier Models Fail on Culturally Specific Data

Why Your Training Data Is Setting You Up to Fail Globally

Frontier models are trained on large corpora of text drawn primarily from the English-language web and Western institutional sources. This is not a secret. What is underappreciated is how deeply that training distribution shapes the model’s outputs, even when it’s being asked to produce content in other languages or for other cultural contexts. The model’s prior, its default assumptions about what is typical, appropriate, or correct, reflects the distribution it learned from. That prior doesn’t disappear when the model switches languages.

Multilingual Capability Won’t Save You From Cultural Failures

One of the most persistent misunderstandings in enterprise AI procurement is treating multilingual capability as a proxy for cultural competence. A model can generate grammatically correct Arabic text while simultaneously encoding assumptions about gender roles, family structure, or political norms that do not reflect the cultural context of Arabic-speaking users. Fluency is a surface property. Cultural alignment is a deeper one.

The distinction matters operationally because evaluation programs built around language capability will miss the cultural alignment failures that determine whether a deployment succeeds or fails in a global market. Model evaluation services that treat cultural alignment as a distinct evaluation dimension, separate from language fluency, surface the failure modes that language-focused benchmarks hide.

The Long Tail of Cultural Knowledge

Cultural knowledge is not evenly distributed across the training data, and the imbalance is not random. High-resource languages with large web presences are well-represented. Low-resource languages and the cultural knowledge embedded in communities that use them are systematically underrepresented. This creates a long tail of failure modes: the model handles high-frequency cultural contexts adequately but fails on the specific cultural knowledge that matters most to underserved user populations.

For enterprises deploying AI in markets where that long tail is the core use case, not an edge case, this is a significant operational risk. The evaluation frameworks designed for high-resource language contexts will not surface those failures because they were not designed to.

Why Your Current Evaluation Program Is Leaving You Exposed

Benchmark Saturation and Its Limits

The most widely used LLM benchmarks now report near-ceiling performance for frontier models. This is sometimes interpreted as evidence that the cultural alignment problem is being solved. It isn’t. It’s evidence that the benchmarks are no longer measuring the right things. Benchmark saturation means the evaluation has stopped differentiating between models on dimensions that matter for global deployment, not that the underlying cultural gaps have been closed.

Research on culturally grounded benchmarks designed to be more challenging than existing leaderboard tests consistently finds that even the best-performing frontier models fall significantly short of human performance on culturally specific knowledge tasks. The gap is not small. It is the difference between a model that appears capable on a benchmark and a model that is actually capable in the deployment context that the benchmark was supposed to represent.

Static Benchmarks Against Evolving Models

Standard benchmarks are also static. Once published, they become part of the training and evaluation ecosystem, which means models can be optimized against them directly or indirectly. A model that scores well on a published cultural benchmark may have been trained on data that overlaps with or was derived from that benchmark. Benchmark contamination reduces the signal value of any static evaluation set over time.

Production-representative evaluation, drawing samples from the actual inputs the model will receive in a specific deployment context, is the evaluation approach that does not suffer from contamination because it reflects what users are actually doing, not what benchmark designers anticipated. Data collection and curation services that source evaluation data from production-like inputs in the target cultural context produce evaluation sets that benchmark contamination cannot undermine.

The Absence of Local Human Judgment

The other thing standard evaluation misses is local human judgment. Evaluating whether a model’s output is culturally appropriate for a specific context requires evaluators who are embedded in that context. An evaluation program that uses Western-trained evaluators to assess outputs for Middle Eastern or Southeast Asian users will miss the specific cultural failure modes that those users will encounter.

This is not a minor calibration issue. The cultural knowledge required to identify certain failures, in moral reasoning, in representation of contested history, in application of local norms to specific scenarios, is not accessible to evaluators who do not share that cultural background. Building evaluation programs around locally embedded human judges is not optional for global deployments. It is what makes the evaluation valid.

What Evaluation Should Look Like

Start With the Deployment Context, Not the Benchmark

Effective cultural evaluation starts with a clear specification of the deployment context: what cultural communities will use the system, what tasks they will use it for, and what cultural knowledge, values, and norms are relevant to those tasks. The evaluation design follows from that specification, not from the availability of existing benchmarks.

This sounds obvious. It isn’t how most enterprise evaluation programs are actually structured. Most evaluation programs start with the available benchmarks and check the model against them. Starting from the deployment context and then designing the evaluation to match it is a different workflow that produces different results.

Culturally Grounded Human Evaluation

The core of a culturally grounded evaluation program is human evaluation by annotators who are embedded in the target cultural context. Those annotators assess model outputs against culturally specific quality criteria: does this response reflect accurate cultural knowledge, apply appropriate norms for this context, and represent contested topics in a way consistent with local perspectives? Model evaluation services that recruit and calibrate evaluators from the specific cultural communities a model will serve produce evaluation programs that are valid for those communities rather than approximations derived from more accessible evaluator populations.

One-Time Evaluations Are a Risk You Can’t Afford

Cultural alignment is not a static property. Models are updated. Deployment contexts evolve. New use cases emerge. An evaluation program that runs once before launch and then stops will miss the drift that occurs as these changes accumulate. Programs that treat cultural evaluation as a continuous operational discipline, running regular evaluation cycles against production inputs and updating the evaluation set as the deployment context evolves, maintain a valid signal of cultural alignment throughout the model’s production life.

How Digital Divide Data Can Help

Digital Divide Data has operated in Cambodia, Laos, Kenya, and the US since 2001, which means our annotator teams are embedded in the cultural communities that global AI deployments are often trying to serve. That depth of local presence is what makes our evaluation and data collection programs culturally valid rather than culturally approximated. 

For programs building culturally grounded evaluation frameworks, model evaluation services design evaluation suites built around the specific cultural context of the deployment, with locally embedded human evaluators who assess outputs against culturally specific quality criteria. For programs building the training data needed to close identified cultural gaps, data collection and curation services, and low-resource languages services source culturally representative training examples from the communities the model needs to serve.

If your evaluation program isn’t measuring cultural alignment for the contexts where you’re deploying, that’s worth addressing before the market tells you about the gap. Talk to an expert.

Conclusion

Frontier models are capable. They are not culturally neutral. The training data that produces their capabilities also shapes their defaults, their values, and their blind spots in ways that systematic standard benchmarks do not surface. For enterprise deployments serving global user populations, that gap is an operational risk that shows up after launch when it could have been identified and addressed before it.

The evaluation programs that find these gaps early share a common structure: they start from the deployment context rather than the available benchmarks, they rely on locally embedded human judgment rather than evaluator populations that don’t share the target cultural background, and they treat evaluation as a continuous discipline rather than a pre-launch gate. The enterprises building this discipline now are not doing it as a compliance exercise. They are doing it because the first mover in a regional market that gets the cultural experience right is the one that earns user trust before a competitor with a less careful evaluation program gets the chance to lose it. That advantage is hard to claw back once a market has decided which provider understands it and which one does not. What’s the gap between what your current evaluation program is measuring and what your deployment context actually requires?

References

Cao, Y., et al. (2023). Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study. arXiv. https://arxiv.org/abs/2303.17466

Li, Y., et al. (2024). CulturalBench: A robust, diverse, and challenging benchmark on measuring the (lack of) cultural knowledge of LLMs. arXiv. https://arxiv.org/abs/2410.02677

Huang, J., & Yang, K. (2023). Culturally aware natural language inference. In Findings of EMNLP 2023. Association for Computational Linguistics. https://aclanthology.org/2023.findings-emnlp.745

Adilazuarda, M. F., et al. (2024). Towards measuring and modeling “culture” in LLMs: A survey. arXiv. https://arxiv.org/abs/2403.15412

Frequently Asked Questions

Q1. Our vendor says their model is already multilingual. Isn’t that enough?

Because standard benchmarks are predominantly designed around Western, English-language contexts. A model can score at the top of a leaderboard while having significant blind spots in the cultural knowledge, values, and norms of non-Western communities. The benchmark was not designed to surface those blind spots, so it doesn’t. Culturally grounded evaluation designed around the specific deployment context is the tool that surfaces them.

Q2. We already ran our own internal evaluation, and the model passed. Why isn’t that sufficient?

Because the team running that evaluation was very likely evaluating against the same kind of benchmark the model was trained to do well on, and very likely did not include evaluators from the specific cultural communities the deployment will actually serve. An internal evaluation that does not include locally embedded judgment from your target markets is not measuring cultural alignment, even if it produced a passing result. The pass tells you the model is technically functional. It does not tell you whether it is culturally appropriate for the markets you are entering.

Q3. This sounds expensive and slow. Can’t we just fix issues as they come up after launch?

You can, but the cost shows up on the other side of the ledger instead. Fixing a cultural misalignment issue after launch means it has already reached real users, generated support escalations, and possibly damaged a regional partnership or a brand reputation you cannot easily rebuild. A culturally grounded evaluation program run before launch is an upfront cost with a defined scope. A post-launch fix is an unplanned cost with a reputational tail attached. Most enterprises that have been through both prefer to pay for the first.

Q4. Our model provider already re-trains and updates the model regularly. Doesn’t that keep cultural alignment current automatically?

On a continuous cadence, not just before launch. Models are updated, deployment contexts evolve, and new use cases emerge. A one-time pre-launch evaluation misses the drift that accumulates as these changes occur. Programs that run regular evaluation cycles against production-representative inputs maintain a valid signal of cultural alignment throughout the model’s production life.

Why Your AI Evaluation Program Is Missing Cultural Failures, and How to Fix It Read Post »

Scroll to Top