Celebrating 25 years of DDD's Excellence and Social Impact.
August 6, 2026 | By Keshav Sharma

Spatio-Temporal Captioning for VLA Model Development

Challenge

A leading autonomous driving company needed multi-modal driving scenarios converted into structured behavioral analyses, combining audio, camera, and telemetry data with scene descriptions, decision reasoning, and safety judgments referenced to regional driving laws. Every decision required human judgment rather than fixed rules, and no annotation guidelines existed at the start, making consistency across a large team the central challenge.

DDD Solution

DDD deployed a 40-member team of full-time labelers and reviewers from its Nairobi delivery center and built the annotation guidelines from scratch through iterative client feedback and daily calibration sessions. Multiple QA checkpoints were introduced, from peer review to a final quality gate, with frame-level precision held to a tolerance of plus or minus one frame. As the project scaled, experienced labelers were promoted internally to reviewer roles, preserving project knowledge and cutting ramp time.

Impact

The team completed a project scope of approximately 36,000 video annotation tasks across five structured batches while maintaining daily throughput targets at every phase. Batch delivery time improved from nine days to four, quality rose well past the agreed acceptance threshold, and the engagement remains active and stable today, with the guidelines now serving as the standard framework for ongoing work.

Scroll to Top