A data flywheel is a self-reinforcing loop: production usage generates data, that data improves the model, the improved model attracts more usage, and more usage generates more data. The idea has been around under different names for years, and in AI right now it is the difference between a model that was good on the day it shipped and a system that keeps getting better without anyone rewriting the architecture. Most teams describe wanting one. Few have actually built the mechanism that makes one turn.
MIT’s research on enterprise GenAI adoption found much the same failure pattern from the other direction: 95 percent of organizations in the study reported no measurable return from their GenAI pilots, and the core barrier the authors identify is learning, since most GenAI systems do not retain feedback, adapt to context, or improve over time.
Meanwhile, Gartner’s tracking of the same problem found that at least half of generative AI projects were abandoned after proof of concept by the end of 2025. The gap between a flywheel that spins and one that does not is rarely the model. It is whether anyone built the loop.
This blog covers what a data flywheel actually requires: the four stages that make the loop close instead of leak, the data engineering underneath each stage, the failure modes that quietly stop a flywheel from turning even when the dashboards look fine, and how to tell a real flywheel from a demo that merely used the word.
Key Takeaways
- A flywheel is a pipeline, not a metaphor. Production usage compounds model quality only when interactions are captured, triaged, labeled, and fed back into training and evaluation as an operating system, not as an occasional cleanup project.
- The loop has four stages, and most programs build one or two. Capture, curation, retraining, and evaluation each have their own failure modes, and skipping any one of them stops the compounding effect even if the other three look healthy.
- Not all feedback is signal. Production interactions include noise, edge cases nobody wants the model to learn from, and outright errors, and a flywheel without a curation stage learns the noise as confidently as the signal.
- Evaluation is what proves the loop is actually compounding. Without a held-out, refreshed evaluation set, a team cannot tell a genuinely improving model from one that has simply drifted toward its own recent outputs.
- A flywheel is only as fast as the labeling behind it. Every stage of the loop, from capture and curation to retraining feedback and evaluation, ultimately runs on human-labeled data, so the speed and quality of that annotation layer set the pace for the whole system.
Why Most GenAI Programs Never Get a Flywheel Spinning
The demo version of a flywheel is easy to describe and easy to fake: point to a chatbot that logs conversations and call the logging a feedback loop. What actually distinguishes a compounding system from a static one is much narrower. It is whether production data that reveals a gap (a wrong answer, a missed edge case, a user correction) gets systematically captured, judged, and turned into a labeled example that changes what the model does next. Most GenAI programs capture the data. Far fewer do anything structured with it. An unused log is not a flywheel. It is an archive.
This is also why the failure mode is so easy to miss internally. A program can have real production traffic, a real logging pipeline, and a real retraining schedule, and still not have a flywheel. None of those things guarantee that the specific interactions worth learning from are being found, labeled, and routed back in. The absence is invisible from the dashboard. It shows up as a model that has been in production for a year and performs almost exactly like it did on day one.
The Four Stages of a Working Flywheel
Stage 1: Capture
Every production interaction that could plausibly reveal a model gap needs to be captured with enough context to judge later: the input, the output, any user correction or explicit feedback, and the surrounding session state. The common failure here is capturing outcomes without context (a thumbs-down with no record of what was asked or why the answer was wrong). That produces a signal nobody can act on months later, once volume makes manual recall impossible.
Stage 2: Curation
Not every captured interaction belongs in the next training run. Curation is the triage step: separating genuine model failures from user error, deduplicating near-identical cases so the loop does not overweight whatever happened to be common that week, and routing ambiguous cases to human review rather than auto-including them. A flywheel without curation does not fail gracefully. It learns its own noise as confidently as it learns real signal. The model quietly gets worse in ways that look like improvement in aggregate metrics.
Stage 3: Retraining and Fine-Tuning
Curated examples feed back into training, whether that is full retraining, targeted fine-tuning, or updates to a retrieval corpus for a RAG system. The operational requirement most programs underbuild here is versioning: which corpus version produced which model version, so that if a retraining cycle makes something worse, the regression can be traced to a specific batch of feedback data rather than triggering a guess-and-rollback cycle across the whole pipeline.
Stage 4: Evaluation
This is the stage that proves the loop is compounding rather than just churning. A held-out evaluation set, refreshed on a cadence and kept separate from anything that entered training, is what lets a team measure whether the retrained model actually improved on the cases that mattered, rather than assuming improvement because the team did work. Without this stage, a flywheel can spin for a year while quietly drifting toward the model’s own recent output patterns rather than toward genuine quality gains, and nobody notices until a metric that was never being tracked finally breaks.
The Curation Problem Is the One Everyone Underestimates
Ask most teams running a flywheel what determines their retraining data, and the honest answer is often volume and recency: whatever came in since the last cycle, roughly deduplicated. That policy sounds plausible and quietly degrades a model over time, because production traffic is not a representative sample of what the model needs to learn. It overrepresents common, easy cases and underrepresents the rare, hard ones that actually move quality. A flywheel that trains proportionally to traffic volume learns to be very good at the easy majority and stays exactly as bad at the consequential minority.
The fix is a curation policy that deliberately overweights informative failures relative to their raw frequency: cases with explicit user corrections, cases flagged by automated quality checks, and a sampled slice of edge cases even when nothing flagged them. That sampled slice matters. It keeps the loop testing itself against the traffic it has not yet mastered, rather than only the traffic it already has.
How Digital Divide Data Can Help
Whether a program builds this loop internally or with a partner, the same four stages decide whether it compounds: disciplined capture, curation that separates signal from noise, versioned feedback into retraining, and evaluation that proves the improvement is real. Producing those is the work we do.
Capture and curation at production scale. Data collection and curation triages production interactions against defined criteria, so retraining data reflects what the model needs to learn rather than just what happened to come through that week.
The labeling layer every stage depends on. AI data preparation turns flagged interactions into structured, versioned training examples, with the provenance record that lets a regression get traced back to a specific batch.
Evaluation that proves the loop is working. Model evaluation services build and refresh the held-out sets that separate genuine compounding quality gains from drift toward the model’s own recent outputs.
If your next question is which of the four stages, capture, curation, retraining feedback, or evaluation, your current loop is actually missing, that’s the assessment we run. Talk to an expert.
Conclusion
A data flywheel is not a property a system acquires by being in production long enough. It is a pipeline someone has to build, with four stages that each have a specific job and a specific failure mode, and the compounding effect only shows up when all four are actually working together. Capture without curation teaches the model noise. Curation without evaluation cannot prove anything improved. Retraining without versioning cannot diagnose a regression when one happens.
The test for whether a program has a real flywheel or a well-instrumented static system: pull the last three retraining cycles and ask what specifically changed in the model’s behavior as a result of each one, measured against a held-out set that predates all three. If that answer is fast and specific, the loop is turning. If it takes weeks to reconstruct, or the honest answer is that nobody measured it, the flywheel is a metaphor the team has been using for a system that is not actually compounding.
References
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-Based Customer Support. (2025). arXiv. https://arxiv.org/abs/2510.06674
Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025). The GenAI divide: State of AI in business 2025 (preliminary findings). MIT Project NANDA. https://nanda.media.mit.edu/ai_report_2025.pdf
Gartner. (2026). Why half of GenAI projects fail: Avoid these 5 common mistakes. https://www.gartner.com/en/articles/genai-project-failure
Frequently Asked Questions
Q1. We retrain our model monthly on new production data. Isn’t that already a flywheel?
A retraining schedule is one stage of a flywheel, not the whole thing, and the missing stages are usually curation and evaluation. Retraining monthly on undifferentiated new data trains proportionally to whatever traffic happened to arrive, which overrepresents easy, common cases and underrepresents the rare, hard ones that actually move quality. And without a held-out evaluation set that predates the new data, there’s no way to confirm the monthly retrain is actually improving anything rather than just moving the model. A useful test: can you show, for your last retrain, exactly what got better and what got worse, measured against a fixed evaluation set? If not, you have a retraining schedule, not a flywheel.
Q2. How much production data do we need before a flywheel starts producing real improvements?
Less than volume alone suggests, because the flywheel’s effectiveness depends on the informativeness of what gets curated in, not the raw count of interactions. A narrow, well-instrumented workflow with disciplined capture and curation can show measurable improvement within a few retraining cycles on a modest volume of genuinely informative examples. A high-volume, uncurated flywheel can run for a year and show nothing, because volume without curation just means more noise at the same signal-to-noise ratio. Start narrow: pick one workflow, instrument it fully, and prove the loop compounds there before scaling the same discipline to more of the product.
Q3. Should the same team that builds the model also curate the flywheel’s feedback data?
They should be closely connected but not solely responsible for the judgment calls, because curation at scale is a distinct discipline from model development, closer to annotation program management than to machine learning engineering. The model team should define what counts as a genuine failure worth learning from and review the curation guidelines, but the actual triage, separating real model gaps from user error, deduplicating, routing ambiguous cases, benefits from the same calibration discipline any annotation program needs: written guidelines, measured agreement, and a defined escalation path. Model teams that try to do this triage themselves alongside their core work tend to under-invest in it precisely because it competes for the same hours as model development.
Q4. What’s the biggest sign that a flywheel has stopped compounding even though it’s still technically running?
Flat or declining performance on a fixed, held-out evaluation set across multiple retraining cycles, while aggregate production metrics look stable or even improve. That divergence is the signature of a flywheel that has started optimizing for its own recent output distribution rather than for genuine quality: production metrics can look fine because the model has gotten good at handling the traffic pattern it already sees a lot of, while the evaluation set, if it is genuinely held out and periodically refreshed with new hard cases, keeps testing the model against what it has not yet learned. A flywheel that only improves on data resembling its own recent history has stopped compounding and started echoing.
Q5. Is a data flywheel worth building for a low-traffic, high-stakes application, or is the volume too low?
Volume matters less than informativeness for high-stakes applications, and these are often exactly the cases where a flywheel earns its cost fastest, because each individual failure carries more consequence and more signal. The adaptation is on the capture and curation side: with lower volume, capture should default to broader inclusion, since there is less traffic to be selective within, and curation should lean more heavily on domain-expert review rather than automated triage, since the cost of learning from a wrongly curated example is proportionally higher. The four-stage structure holds regardless of volume; what changes is how much automation versus expert judgment each stage uses.

Kevin Sahotsky leads strategic partnerships and go-to-market strategy at Digital Divide Data, with deep experience in AI data services and annotation for physical AI, autonomy programs, and Generative AI use cases. He works with enterprise teams navigating the operational complexity of production AI, helping them connect the right data strategy to real model performance. At DDD, Kevin focuses on bridging what organizations need from their AI data operations with the delivery capability, domain expertise, and quality infrastructure to make it happen.