OpenAI relies on a workforce of around 10,000 contractors to rate and improve its AI models, and several of those contractors have been fired for using AI to do the training work, according to a 404 Media investigation. The scale and structure of the program mirrors the quality rater systems used by Google to test search and AI results, but the method of improvement and the risk it introduces are different: models tuned on AI-generated feedback can degrade over time, a problem known as model collapse.
What OpenAI’s quality rater and AI trainer program does
OpenAI hires thousands of contractors whose job is to evaluate model outputs, score responses, and provide feedback that shapes future versions of the same models. The work is similar in purpose to Google’s quality rater program, which uses large human panels to rate search results and AI Overviews so that ranking systems can be tuned against human judgment.
The 404 Media report found that OpenAI’s contractor base runs into the thousands and that multiple contractors were terminated for completing training tasks with AI tools rather than by hand. The contradiction is sharp: the people hired to improve an AI model were using an AI model to do the improving, which undermines the human signal the program is meant to collect.
Why model collapse is the real concern
Model collapse is what happens when AI models are repeatedly trained on text or ratings produced by previous AI outputs, which is precisely the risk when contractors use AI to complete training tasks intended to capture human feedback. Each subsequent round of training on AI-generated material narrows the variety of patterns the model sees, degrades performance on edge cases, and pushes outputs toward patterns that match prior AI text rather than human writing.
Researchers studying the phenomenon have warned that if a meaningful share of the rater workforce uses AI to generate their feedback, the feedback itself becomes AI-flavored. Models tuned on that feedback then drift further from the kinds of human writing and judgment the feedback was supposed to represent. Catching and removing the offending contractors is the immediate fix; preventing the same pattern at scale is the harder problem.
How does OpenAI’s program compare to Google’s quality rater program?
Google has run a quality rater program with thousands of human reviewers for years, and the reviewers rate search results against published guidelines. The program is also used to test AI features in search. OpenAI’s AI trainer program operates on a similar scale, with a reported headcount around 10,000, and serves a parallel purpose: giving the lab human judgments to measure and tune its models against.
The core difference is what happens when rater output is contaminated. For search, contaminated rater scores degrade ranking quality, which Google can detect through live result monitoring and ranking volatility tracking. For generative models, contaminated scores feed the next training run directly, which is why contractors caught using AI to complete rating tasks were fired.
What contractors were actually doing
According to 404 Media, the fired contractors used AI tools to generate, evaluate, or rewrite the content they were paid to produce by hand. The work itself ranges from rating model responses and flagging bad outputs to writing examples the model is expected to learn from. Any of these tasks, when completed with an AI assistant, pushes AI-generated material back into the training pipeline.
OpenAI has not published a breakdown of which tasks were most affected. According to 404 Media, the fired contractors were spread across multiple contractors and were not limited to a single product line, which suggests the risk was systemic rather than confined to one team.
Why this matters for anyone building with AI
Two practical lessons follow from the report. First, post-training data quality is now a known failure mode, and labs that scale human feedback programs will need automated detection on top of policy enforcement. Second, anyone using public model outputs to fine-tune or evaluate their own systems should expect the same drift if the source material is itself AI-generated.
For businesses shipping AI features, the safe defaults are clear: keep a human-in-the-loop on rating work, audit a sample of rater outputs for signs of AI assistance, and treat any dataset that mixes human and AI writing as a separate category when measuring model quality.
FAQ
Does OpenAI have quality raters like Google?
Yes. OpenAI runs a contractor-based AI trainer program with a workforce reported in the thousands, serving the same function as Google’s quality rater program: providing human judgment to evaluate and tune model outputs.
Why were OpenAI AI trainers fired?
Several contractors were fired for using AI tools to complete the training and rating tasks they were hired to do. The terminations were confirmed in a 404 Media report covering multiple contractors across the program.
What is model collapse?
Model collapse is the degradation that happens when AI models are trained on AI-generated data, causing the variety of patterns to narrow and outputs to become less reliable across further training rounds.
Try the AI visibility report
The AI visibility report runs a full technical audit of a site and shows the measured result behind every check. Open the AI visibility report.
This article summarizes reporting from seroundtable.com.


