A model failure is not always a model problem.
Issues found during road testing became engineering tickets. An engineer would adjust the model and test again—but some failures needed better or more representative data, not another round of parameter tuning.
There was no scenario-level record connecting the issue to the data collected for it, the preparation work requested, or the model results that followed. Teams could track a ticket. They could not track the improvement cycle.
I followed the data across teams—not just screens.
I interviewed model engineers, data specialists, engineering leadership, annotation operations, and data publishing teams. Journey mapping exposed where context disappeared as data moved between people and systems.
The design opportunity was to make the failure scenario a persistent unit of work: a shared thread connecting the data assembled for it, operational requests, released datasets, and evaluation evidence.
Tight collaboration made the service real.
The product was shaped alongside the data specialists who would operate it. We reviewed ideas as they emerged—from dataset curation to comparing reference labels with model output—and iterated before the workflows hardened into infrastructure.
During an onsite hackathon, the team brought one collaboration feature from idea to a working flow in a week. A data specialist could share curated samples with a model engineer, receive flagged examples asynchronously, refine the data-mining query, and then send the improved dataset for managed annotation.
Useful first. Integrated second.
Many capabilities began as lightweight spreadsheet or notebook workflows. These prototypes let the team test whether a workflow was useful before investing in the APIs and large-scale data operations needed to support it in the shared interface.
This approach also clarified product boundaries. The workspace needed to support operational curation, while deeper exploratory analysis was often better served by dedicated analytics tools.
One workspace for the full improvement loop.
The final experience brought curation, operational requests, and evaluation into one scenario-centered workflow. The highlights below show three of the more useful interface capabilities without trying to replace every specialized system around them.
Data specialists could tailor the metadata shown for a review task, making large datasets easier to scan without forcing every workflow into the same display.
Simplified recreation shown to protect proprietary information.
Reviewers could move from KPI trend analysis to the worst-performing examples, then inspect a single example in context.
Simplified recreation shown to protect proprietary information.
Teams could submit and track labeling or data-collection requests from the same workspace.
Simplified recreation shown to protect proprietary information.
Scenario tracking became a shared foundation.
By the end of the project, the organization had moved from no scenario-centered tracking to a central workspace spanning dozens of models and thousands of scenarios.
Connected systems
Collection, annotation, dataset publishing, and evaluation gained shared connection points.
Self-service visibility
Model engineers could observe data work and follow progress without relying on disconnected status updates.
Persistent tracking
Teams could follow a failure scenario across repeated data and model iterations.
The workspace became the de facto scenario-centered data-management tool.
The interface was only the visible layer.
My work ranged from journey mapping and information architecture to coded prototypes, interview plans, engineering handoff, and UI quality assurance. The harder design problem was aligning services, teams, and feedback loops around the same unit of work.
We deliberately deferred synthetic-data supplementation until the core lifecycle had been proven with collected data. Establishing reliable tracking first created a foundation for a more automated improvement loop later.
The system’s value is also reflected in what came next. I’m now building on these principles for related model-development use cases beyond the original domain, extending an approach the organization recognized as broadly useful.