Celebrating 25 years of DDD's Excellence and Social Impact.
TABLE OF CONTENTS
    AI Data budget

    How to Set a Realistic AI Data Budget: What Programs Actually Spend vs. What They Plan

    Kevin Sahotsky

    There’s a specific moment in AI program planning where budgets go wrong, and it isn’t the estimate. It’s the line items. The plan has a model line, a compute line, an integration line, and maybe a tooling line. Then, six months in, the actual spend starts accumulating in categories the plan never named: annotation that senior engineers were quietly doing themselves, a second pass of labeling after the guidelines changed, an evaluation set that had to be built from scratch because nobody budgeted one, and rework on a dataset that looked cheap until the quality numbers came back.

    This isn’t a niche problem. In its March 2025 forecast, Gartner put worldwide GenAI spending at 644 billion dollars for 2025, an increase of 76.4 percent from 2024. Its July 2024 press release put GenAI deployment costs at $5 million to $20 million, depending on the approach, and named escalating costs among the top reasons projects are abandoned after proof of concept, right after poor data quality. Its 2026 follow-up found the outcome was worse: at least half were abandoned. Gartner’s guidance on GenAI total cost of ownership is blunt about the pattern: total costs often exceed initial expectations because of hidden items like compliance reviews, model retraining, and internal overheads. Data operations are embedded in almost every one of those hidden items.

    This blog is about closing the gap between the budget you plan and the budget you’ll actually spend. It covers the line items programs consistently omit, the cost drivers that actually move data spend, a construction method that works backward from model requirements instead of forward from a per-label price, and the trade you should make when the number comes back too high.

    Key Takeaways

    • Budgets fail by omission, not underestimation. The model and compute lines are usually close; the categories that blow up plans are the ones that never appeared: ongoing annotation, rework, evaluation sets, edge case collection, and guideline development.
    • Data work is an operating cost wearing a project cost’s clothes. Programs budget data as a one-time acquisition and then discover that retraining, drift response, and production feedback all consume labeled data continuously.
    • The pilot hides the real number. Pilot-phase data costs are invisibly subsidized by senior staff doing annotation themselves at a scale where that’s possible, which makes the production estimate look inflated when it’s actually the first honest number.
    • Per-label price is the least informative number in the budget. Cost per accepted, quality-verified label, including rework and QA, is the number that predicts what you’ll spend; the cheapest per-label quote is frequently the most expensive dataset.
    • When the budget is fixed, cut volume before quality. Volume can be added back cleanly when budget returns; a degraded quality tier and a missing evaluation set cannot be cheaply repaired.

    Where Planned and Actual Budgets Diverge

    The Lines Programs Plan

    A typical AI program budget names the visible categories: model development or licensing, compute and inference, integration engineering, tooling, and sometimes an initial dataset purchase or annotation project. These estimates are usually defensible. Teams benchmark compute, vendors, quote integration, and the initial dataset gets a per-label quote that looks precise.

    The Lines Programs Discover

    The actual spend accumulates elsewhere. Guideline development and calibration: the unglamorous work of turning model requirements into instructions annotators can apply consistently, including the pilot rounds where inter-annotator agreement gets measured and the guidelines get revised. Rework: the second and third passes that follow every guideline change, every edge case discovery, and every QA finding, in a program where the first pass was priced as if it were the only pass. 

    Evaluation sets: the human-verified gold data that quality measurement and drift monitoring depend on, which almost no first budget contains because it doesn’t feel like training data. Edge case collection: the deliberate sourcing of the rare cases production will surface, which is a data program of its own. And the production loop: the continuous annotation of production failures that separates improving systems from stalling ones. Across the program budgets I’ve reviewed, it’s common for these unnamed categories to end up rivaling the initial dataset line itself, not because any one of them is large but because all of them recur.

    The whole argument fits in two columns:

    Lines programs plan Lines programs discover
    Model development or licensing Guideline development and calibration rounds
    Compute and inference Rework passes after guideline and edge case changes
    Integration engineering Evaluation set construction and maintenance
    Tooling and platforms Deliberate edge case collection
    Initial dataset (one-time, per-label quote) The production loop: continuous annotation of production failures, drift response, refresh cycles

    The left column is priced in every plan. The right column is where the overruns live, and every item in it recurs.

    The Five Drivers That Actually Move Data Spend

    Task ambiguity is the first driver, and the least priced-in. Labeling a stop sign and grading the helpfulness of a model response are both ‘annotation,’ but the second requires judgment, calibration, and adjudication of disagreements. All of that is time. The more ambiguous the task, the more the real cost sits in guideline quality and calibration rather than in the labeling itself.

    Quality tier is the second. The QA design that supports a demo differs from one that supports a regulated deployment: sampling rates, review tiers, agreement thresholds, and documentation all scale with the consequence of being wrong. Budgeting quality as a percentage bolt-on misses that quality is a design choice with its own cost curve.

    Domain expertise is the third. Generalist annotation and specialist annotation, clinicians, lawyers, robotics-literate reviewers, occupy different labor markets. If the task needs the specialist, the budget either pays for it or pays more later in rework.

    Volume dynamics are the fourth. Data needs don’t arrive flat. They spike at retraining, at expansion into new domains, and after every drift event. A budget built on average monthly volume will be wrong in both directions: idle capacity in quiet months, missed deadlines in spikes. What you’re actually buying is capacity with a ramp profile, and it should be priced that way.

    Change is the fifth, and the most reliably omitted. Guidelines evolve as the model and the product evolve. Every material guideline change ripples into re-annotation of affected data. Programs that budget zero for change are implicitly assuming the first guidelines will never need revision.

    A Construction Method That Produces a Defensible Number

    Start from the model requirements, not the label price. What does the model need to learn, at what quality, refreshed how often? That converts into annotation volume with a quality tier and a cadence. Price the unit honestly: cost per accepted label, meaning the all-in figure that includes QA, adjudication, and expected rework, not the raw per-label quote. Then split the budget into build and run. The build phase covers the initial corpus, guideline development, calibration, and the evaluation set. The run phase covers the ongoing loop: production sampling, edge case annotation, drift response, and refresh cycles. In my experience, teams that present data as build-plus-run get their budgets approved more often than teams that present a single dataset number, because finance recognizes the shape: it looks like an operating capability, which is what it is.

    Two sanity checks before the number goes in the deck. First, the evaluation set has its own line, because unbudgeted quality measurement rarely gets built. Second, rework carries an explicit allowance; a first-pass-only budget is a bet that your first guidelines are your final guidelines, and nobody has ever won that bet.

    When the Number Comes Back Too High

    The wrong response is to shop the per-label price down until the number fits, because the quote that undercuts the market is usually recovering its margin from your rework budget. The right response is to descope volume while protecting quality: a smaller, well-covered, quality-verified dataset with a real evaluation set beats a large degraded one, and it leaves a foundation that scales cleanly when more budget arrives. Descope the corpus, keep the QA design, keep the eval set, keep the calibration. Those are the parts you can’t cheaply add back later.

    How Digital Divide Data Can Help

    The build-and-run anatomy above is exactly what we construct with clients, so here’s how it maps to a real engagement.

    A quote you can put in a budget deck. After seeing your data, we price against your quality tier and task ambiguity with QA and expected rework, factoring them into the number, whether the scope is sourcing and curating new datasets or preparing and labeling the data you already hold. Cost per accepted label is the figure you plan on.

    The lines plans forget, delivered as line items. Guideline development, calibration rounds, and the evaluation sets we build and maintain for quality and drift measurement, each explicitly scoped and priced rather than surfacing as overruns.

    Run-phase capacity with a ramp profile. Throughput commitments that flex with retraining cycles and drift response, with the pipelines that route production data back into annotation and training built alongside, so the loop is a budgeted operation rather than a surprise.

    If you’re building next year’s AI budget now, a scoping conversation before the number is locked costs nothing and tends to save the change orders. Talk to an expert.

    Conclusion

    The gap between planned and actual AI data spend isn’t an estimation error. It’s a categories error: the plan prices the visible dataset and omits the operating loop that production actually runs on. The fix is structural, not heroic. Name the hidden lines, price the accepted label rather than the raw one, split build from run, protect the evaluation set, carry a rework allowance, and when the total is too high, cut volume before quality.

    Here’s the one-line test for your current plan: does the data budget survive contact with the second version of your annotation guidelines? If a guideline revision would blow the number, the number was never realistic. It was just early.

    References

    Gartner. (2025, March 31). Gartner forecasts worldwide GenAI spending to reach $644 billion in 2025. https://www.gartner.com/en/newsroom/press-releases/2025-03-31-gartner-forecasts-worldwide-genai-spending-to-reach-644-billion-in-2025

    Gartner. (2024, July 29). Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025. https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025

    Gartner. (n.d.). Enterprise guide to generative AI: Expert insights on ROI, use cases, and cost management. Accessed August 2026. https://www.gartner.com/en/topics/generative-ai

    Gartner. (2026). Why half of GenAI projects fail: Avoid these 5 common mistakes. https://www.gartner.com/en/articles/genai-project-failure

    Frequently Asked Questions

    Q1. What does annotation actually cost per label? Give me a number.

    Any number quoted before seeing your task is a marketing number, and that’s the honest answer. The same ‘label’ spans an order of magnitude of cost depending on task ambiguity, quality tier, domain expertise, and rework expectations, which is why the useful question is different: what is the cost per accepted label at my quality bar, all-in? Get that figure quoted against a sample of your real data with your real guidelines, including QA and an explicit rework assumption. Two vendors quoting the same raw per-label price can differ materially on that all-in figure, and the all-in figure is the one your budget will actually experience.

    Q2. Should AI data be budgeted as a project cost or an operating cost?

    Both, explicitly split. The build phase, the initial corpus, guideline development, calibration, and the first evaluation set, behave like a project cost with an end date. The run phase, production sampling, edge case annotation, drift response, refresh cycles, and evaluation set maintenance, is an operating cost that persists as long as the model serves traffic. Programs that budget only the build phase rediscover the run phase as overruns; programs that present both get cleaner approvals because the structure matches how finance already thinks about capabilities versus purchases.

    Q3. How do I justify budget for evaluation sets when they don’t train the model?

    Frame them as the instrumentation, because that’s what they are. Without a maintained, human-verified evaluation set, the program cannot measure production quality, cannot detect drift before the business metric moves, and cannot prove that any retraining actually improved anything, which means every other dollar in the budget is spent unmeasured. The evaluation set is typically a small fraction of total data spend, and it is the fraction that makes the rest auditable. If a stakeholder wants it cut, the counter-question is direct: which of our quality claims are we comfortable making without evidence?

    Q4. Our pilot data costs were low. Why is the production quote so much higher?

    Because the pilot number wasn’t a cost, it was a subsidy. Pilot data is typically hand-assembled and hand-labeled by senior engineers and data scientists whose time was charged to salaries rather than to the data line, at a volume where that’s feasible. Production removes the subsidy: volume exceeds what senior staff can absorb, quality needs formal QA rather than familiarity, and edge cases need deliberate sourcing. The production quote is not inflated; the pilot cost was artificially low. The useful comparison is the production quote against the fully loaded cost of your engineers doing the same work, which is a comparison the quote usually wins.

    Q5. The budget is fixed, and the data estimate exceeds it. What do we cut?

    Cut volume, protect structure. Reduce the corpus size and narrow the initial domain coverage, but keep the quality tier, the calibration process, the evaluation set, and the rework allowance intact. A smaller dataset at verified quality produces a better model and a truthful measurement of it, and it scales cleanly when budget returns. The tempting alternative, keeping the volume and dropping the quality tier or the eval set, produces a larger dataset you can’t trust and a model you can’t measure. Repairing both later reliably costs more than the savings.

    Get the Latest in Machine Learning & AI

    Sign up for our newsletter to access thought leadership, data training experiences, and updates in Deep Learning, OCR, NLP, Computer Vision, and other cutting-edge AI technologies.

    Scroll to Top