Every AI data partner you talk to will tell you they have high-quality, deep expertise, and flexible pricing. Every deck looks the same. Every reference call goes well, because nobody offers you the reference that went badly. And yet the outcomes across this market are wildly uneven: some teams get a partner who quietly compounds their model quality over years, and some get eighteen months of rework, missed deadlines, and labels they end up redoing in-house.
I lead strategic partnerships and go-to-market at Digital Divide Data, which makes me an interested party. Every item on this checklist is independently verifiable, which is the only reason a vendor-written version of it is worth reading.
In a 2026 analysis, Gartner found that at least half of GenAI projects were abandoned after proof of concept by the end of 2025, worse than the 30 percent it had projected in its original 2024 forecast. Gartner attributes the abandonment to poor data quality, inadequate risk controls, escalating costs, and unclear business value. Of those four, one is largely determined before the project starts, by a decision most teams treat as procurement: who prepares your data.
Key Takeaways
- Evaluate the operation, not the pitch: Look for clear evidence of quality through sampling methods, agreement scores, escalation paths, and calibration processes.
- Test domain expertise directly: Ask the annotation team to work through real edge cases from your data to assess their practical understanding.
- Treat the pilot as the real evaluation: A paid pilot with agreed metrics provides a clearer view of performance than references or sales claims.
- Assess workforce stability: Low attrition and strong team continuity are critical for maintaining consistent annotation quality over time.
- Look beyond low per-label pricing: Lower upfront costs can quickly be offset by rework, relabeling, QA issues, and additional engineering effort.
Why This Decision Carries More Weight Than It Looks Like It Does
A data partner isn’t a supplier in the normal sense. A supplier who ships a bad batch of components costs you that batch. A data partner who ships subtly inconsistent labels costs you a training run, then the debugging cycle where your engineers assume the model is the problem, then the discovery, then the re-annotation, then the retraining. The failure is expensive precisely because it’s slow to surface: bad labels don’t announce themselves; they just quietly cap your model’s ceiling.
A pattern worth naming concretely, without identifying details: a computer vision program hit a quality plateau that survived two model architecture changes and a full retraining cycle. Engineering spent six weeks debugging the model before anyone re-audited the training labels and found that annotators disagreed on roughly 15 percent of a rare-class category, not because the class was hard to see, but because the original guideline never resolved an edge case that kept coming up. Relabeling that one category, without touching the model at all, moved the metric more than either architecture change had. The plateau had been treated as a model problem for the better part of a quarter. It was a label problem.
That asymmetry is why the evaluation deserves more rigor than most procurement processes give it. The good news is that the signals that predict a strong partner are observable during evaluation, if you know where to look. Here’s where to look.
The Seven Things to Actually Evaluate
- QA Methodology They Can Show, Not Describe
Every vendor says they have rigorous QA. The question is whether they can show you the machinery. Ask for the sampling design on a live program: what percentage of output gets reviewed, how the review tiers are structured, what triggers escalation. Ask for inter-annotator agreement numbers from a real project in a domain adjacent to yours, and ask how those numbers are measured and how often. A partner with a real QA operation answers these in specifics within a day. A partner who responds with adjectives usually has not built one.
- Domain Expertise You Can Test in an Hour
Generic annotation capacity and domain-trained teams look identical in a deck and completely different on your data. The fastest test I know: pull three genuinely ambiguous examples from your own dataset, the edge cases your internal team debates, and ask to walk through them with the people who would actually run your program, not the sales engineer. How they reason about ambiguity, whether they ask the right clarifying questions, and whether they’ve seen your failure modes before tells you more than any case study.
- Guideline Development as a Collaboration, Not a Handoff
Annotation guidelines are where model requirements become label behavior, and the partners who produce great data treat guideline development as joint work: they push back on ambiguous instructions, propose edge case handling you hadn’t considered, and run calibration rounds before production. Partners who accept your first-draft guideline without questions aren’t being easy to work with. They’re skipping the step where most label quality is actually determined.
- Security and Compliance That Matches Your Exposure
The certifications that matter depend on your data. If you’re handling health data, HIPAA compliance isn’t optional. If you’re operating in Europe, GDPR (the EU’s General Data Protection Regulation) applies. ISO 27001 and SOC 2 are the baseline signals that security practices are audited rather than asserted. Beyond the certificates, ask operational questions: where does the data physically reside, who can access it, and what happens to it when the engagement ends. Certificates alone do not answer those questions.
- Workforce Model and Attrition
This is the evaluation criterion buyers skip most often and regret most often. Annotation quality lives in calibration, and calibration lives in people. Every annotator who leaves takes months of accumulated task understanding with them, and their replacement starts the learning curve over, on your budget. Ask for attrition rates directly. Ask whether the team assigned to your program stays with your program. A partner whose workforce model is built for continuity will answer proudly; a partner running a churn model will answer vaguely.
- Scalability With Commitments, Not Aspirations
Your volume will spike, your deadlines will compress, and the question is what happens then. Ask for throughput commitments in writing: ramp time to add capacity, turnaround at your peak volume, and quality guarantees that hold during ramps. The critical follow-up is how quality is protected while scaling, because adding annotators is easy and adding calibrated annotators is not. A real answer describes the onboarding and calibration pipeline for new team members. An aspirational answer offers no such description.
- Pricing Structure That Doesn’t Fight Your Interests
Pure per-label pricing creates an incentive to maximize throughput, and throughput pressure is where quality quietly dies. That doesn’t make per-unit pricing wrong, but it makes the question worth asking: what in the commercial structure rewards accuracy rather than volume? Quality-linked terms, rework provisions that put the cost of bad labels on the vendor, and pilot pricing that isn’t a loss-leader teaser all signal a partner planning to win on quality rather than on lock-in.
Red Flags That Predict the Bad Ending
A few patterns show up disproportionately in the engagements that go wrong. A vendor who quotes a firm price before seeing your data is pricing a fantasy, and the correction will arrive as change orders. A vendor who won’t put quality metrics in the contract is keeping quality as a discussion topic rather than an obligation. A vendor who can’t introduce you to the delivery team before signing is selling you a team that doesn’t exist yet. And a vendor whose answer to every capability question is yes has stopped evaluating fit and started closing. None of these is disqualifying alone. Two together should slow you down. Three should end the conversation.
The Pilot Is the Real Evaluation
Everything above narrows the field. The pilot decides it. A well-designed pilot is paid, because free pilots get the vendor’s spare capacity rather than their real operation. It runs on your data, including a deliberate slice of your edge cases, not a curated sample. And its success metrics are agreed in writing before it starts: target accuracy against a gold set you control, inter-annotator agreement thresholds, turnaround times, and the guideline iteration process. In my experience, two to four weeks of pilot at meaningful volume surfaces the operational truth that six months of sales conversations cannot. The vendors worth hiring welcome this structure, because it’s the arena where a real operation beats a good deck.
How Digital Divide Data Can Help
So how do we score against our own list?
QA you can inspect: Our programs run tiered review with inter-annotator agreement measured continuously, and we share the numbers, sampling designs, and escalation paths from comparable programs during evaluation, not after signing.
Teams that stay: Our workforce model is built around continuity: the team that calibrates on your program stays on your program, which is why low attrition is one of the things clients cite most when they renew.
Security that’s audited: ISO 27001 certification and SOC 2 Type II attestation, plus GDPR and HIPAA compliance programs, with operational answers about data residency, access control, and what happens to your data when the engagement ends.
A pilot on your terms: your data, your edge cases, and metrics agreed in writing before it starts. We run these across data collection and curation, AI data preparation, and model evaluation.
Bring us your seven-point checklist. We’ll answer it in specifics, starting with a pilot on your data. Talk to an expert.
Conclusion
The AI data partner decision is unusual: the failure mode is slow, expensive, and disguised as a model problem, and the marketing across the market is indistinguishable. What separates those two outcomes is not luck. It is whether the buyer demanded evidence instead of assurance, and whether a paid pilot got the final word before the contract did.
One last suggestion: write your evaluation criteria down before the first vendor call, not after. Criteria formed during the sales process have a way of drifting toward whatever the most polished pitch happened to emphasize. What’s actually on your list right now, and how many of the seven above are on it?
Frequently Asked Questions
Q1. Isn’t a vendor writing a vendor-evaluation guide a conflict of interest?
Yes, and it’s better to name it than to pretend otherwise, which is why my role is stated in the second paragraph. The mitigation is that everything in this checklist is verifiable independently: IAA numbers, attrition rates, certifications, pilot metrics, and contract terms are facts you check, not claims you take from me. A biased checklist made of checkable items is still a useful checklist. And commercially, quality-focused vendors benefit from educated buyers, because uneducated buyers select on price and polish, which is exactly the selection process that burns them.
Q2. We already have an internal labeling team. Do these criteria still apply?
Most of them, yes, and running your internal team through the same checklist is clarifying. Internal teams often score well on domain expertise and security and surprisingly poorly on QA methodology, throughput commitments, and calibration processes, because those disciplines were never formalized. The build-versus-partner question usually resolves into a hybrid: internal teams own guidelines, gold sets, and final judgment, while a partner provides calibrated capacity and QA infrastructure. The checklist tells you which pieces you actually have.
Q3. How much should we expect to pay for a pilot, and what if the vendor offers it free?
Expect to pay something meaningful relative to the work performed, because you want the vendor’s production operation, not their spare capacity. A free pilot isn’t disqualifying, but it changes what you’re measuring: free pilots are often staffed by the best available people as a sales investment, which tells you the vendor’s ceiling rather than their standard delivery. If you accept a free pilot, compensate by insisting on the same structure you’d demand from a paid one: your data, your edge cases, metrics agreed in writing, and an explicit statement of whether the pilot team is the delivery team.
Q4. What’s a reasonable inter-annotator agreement number to require?
It depends on task ambiguity, which is why demanding a universal number is the wrong move and demanding the measurement is the right one. In our experience, well-calibrated teams on well-specified tasks commonly sustain agreement in the 85 to 95 percent range, while genuinely ambiguous judgment tasks can sit lower without indicating a problem. What you should require: agreement measured continuously rather than once, reported at the subgroup and category level rather than only in aggregate, and a defined process for what happens when it drops. A vendor comfortable with that requirement has a real quality operation.
Q5. How long should we expect vendor evaluation to take, and can we shorten it?
A serious evaluation with a properly structured pilot typically runs eight to twelve weeks end to end: two to three weeks for the paper evaluation and team interviews, two to four weeks of pilot, and the remainder for metric review and commercial negotiation. You can compress the paper phase substantially by sending your checklist and edge cases before the first call and disqualifying on the responses. You should not compress the pilot, because the pilot is the only phase producing evidence rather than claims. Teams under deadline pressure sometimes skip it and select on references and price; that decision is exactly how buyers end up getting burned.

Kevin Sahotsky leads strategic partnerships and go-to-market strategy at Digital Divide Data, with deep experience in AI data services and annotation for physical AI, autonomy programs, and Generative AI use cases. He works with enterprise teams navigating the operational complexity of production AI, helping them connect the right data strategy to real model performance. At DDD, Kevin focuses on bridging what organizations need from their AI data operations with the delivery capability, domain expertise, and quality infrastructure to make it happen.