OpenAI Raters Are Being Cut for Using AI to Grade AI
Contractors hired through Mercor to rate ChatGPT's answers are being removed for using chatbots, Grammarly and translation tools, and the rules show how hard it is to keep human feedback human.
Reinforcement learning from human feedback has a simple premise. A person reads what the model wrote and judges it. The model learns from that judgment. If the person quietly hands the judging to another model, the premise breaks, and the lab pays for human feedback it did not get.
That is the problem at the center of a 404 Media report published September 22 by Joseph Cox. Contractors working on OpenAI training projects are being removed for using AI to do the work, including large language models, the AI detector GPTZero, Grammarly and AI translation services.
Who the contractors work for
Precision matters here, because some of the coverage that followed got it wrong. The contractors 404 Media spoke to were not OpenAI employees. Two of them worked through Mercor, the AI training-data company that hires experts for lab projects. In its earlier reporting on OpenAI's Project Lily, 404 Media said contractors were recruited by Crossing Hurdles and paid through Mercor. OpenAI sets the project and the rules, but the people are engaged and removed through a vendor chain.
Mercor confirmed the policy on the record. "When we confirm an expert has used AI to complete a task, we immediately remove them from the project," a spokesperson told 404 Media. OpenAI declined to comment.
The 10,000 number, read correctly
Some headlines turned this story into "OpenAI fires 10,000 contractors." That is not what the reporting says. According to 404 Media, one internal document states that these projects "can include more than ten thousand contractors." That is the size of the workforce, not the number removed.
No one has published a count of removals. The closest thing to a number is a contractor's remark to 404 Media that in a group of thousands, "tons" have been caught. Another called AI use the one thing that gets you removed fastest. 404 Media spoke with three contractors doing OpenAI work, plus a fourth who trained models for several companies. That is a small sample, and the scale of the problem remains unknown.
What the rules say
The internal guidelines are blunt. Contractors may not use LLMs, GPTZero or other AI detectors, Grammarly, or AI translation on annotation tasks, and the stated penalty is immediate removal from the project. The ban on detectors cuts both ways: reviewers checking contractors' work are also told not to run their output through detection tools. One document says GPTZero and similar tools "are not reliable."
Instead, detection is done by people. Reviewers look for repetitive wording, heavy use of em dashes and work completed faster than anyone could have read the material. The guidelines also tell reviewers not to explain to contractors why they suspect AI, because it would be easier for them to hide if they knew what reviewers were looking for.
That is a telling choice. OpenAI builds some of the most capable language models in the world, and its own contractor rules say automated AI detection is not good enough to base a decision on. Enforcement rests on the judgment of other humans, the same kind of judgment the project is trying to buy in the first place.
Why the work matters
The stakes are clearer once you know what the work is. On September 14, 404 Media reported on Project Lily, in which hundreds of contractors read real ChatGPT users' prompts and conversations, then rate and critique the model's responses. IBTimes, citing that report, said reviewers score candidate responses on a 1 to 7 scale and that one contractor described pay of more than $50 an hour. Among the goals is training ChatGPT to be less sycophantic and to avoid anthropomorphizing itself.
That is subtle judgment. Deciding whether an answer is quietly flattering the user, or whether a model is presenting itself as more human than it is, is exactly the kind of call a model would get wrong about itself. A contractor who pastes the conversation into a chatbot and asks for a rating is feeding the model's own blind spots back into its training signal.
The Project Lily reporting also raised a separate privacy issue. OpenAI told 404 Media it uses a filter to strip identifying details before human review, but acknowledged it can miss uncommon identifying information. IBTimes reported that eligibility for this review is on by default for Free, Plus and Pro consumer accounts and off by default for Enterprise, Business and Education. Contractors using outside AI tools would add a second layer of exposure, sending those user excerpts to third-party services.
The incentive problem
None of this is surprising from the contractor's side. Rating work is often paid per task or tracked on speed, and AI tools make it faster. One fired contractor told 404 Media they needed a little boost and turned to AI, which led to their removal. When the pay structure rewards throughput and the tool that raises throughput is a click away, some share of a workforce in the thousands will use it.
For the labs, this is a supply-chain problem, not a discipline problem. Human feedback now flows through several layers: the lab, a data company like Mercor, sometimes a recruiter, then the individual expert. Each layer has its own incentives, and the lab sees the work only at the end. Firing people after the fact catches some of it. It does not tell the lab how much AI-written feedback already reached training.
What to watch
This lands at an awkward moment for the human-data business. On the same day 404 Media published, Snorkel AI and micro1 announced new valuations of $3.5 billion and $4 billion on the premise that expert human judgment is scarce and worth paying for. Mercor, by TechCrunch's count, reached $2 billion in gross annualized revenue this summer.
That premise now carries a verification cost. Expect data vendors to compete on proof of human work as much as on expert credentials: tighter time tracking, stricter tooling, more review layers. Expect labs to ask for it in contracts. The question 404 Media's report raises is not whether some raters cheat. It is whether anyone in the chain can show, at scale, that the human feedback they sold was actually human.
