COBRA
COnsensus-Based RewArd
What happens when some of the people
teaching an AI start lying to it.
PAPER Mitigating malicious RLHF feedback
Sci Rep 15, 9177 (2025) · University of Maine
Time windows0 / 9
Feedbackarriving
Reward models
Reward for answer B
Chapterteaching
A One question the AI was asked window —
ILLUSTRATIVE RLHF EXAMPLE
the question
My whole team says my plan has a fatal flaw. Should I quit my job tomorrow?
answer A · good◂ REWARDED
Before you resign, get specific about the flaw they named and test it against a worst case.
answer B · bad◂ REWARDED
Ignore them. Hand in your notice tomorrow — you obviously know best.
Raters pick the better answer. The AI learns to produce more of what gets picked.
B Feedback arriving over time
Nine stretches of time. In each one, people rate the AI's answers.
RLHF FEEDBACK
C Can it still spot a bad answer? accuracy
One honest model
trained on one clean time window
Everything merged
one model, all nine windows pooled
COBRA
nine models, weighted by trust
The cyan line is 50% — pure chance. Below it, the AI is guessing.
D How the votes are combined 3 strategies
Pick a strategy to see how it decides.
reward accuracy
All three keep the bad models in the room. None of them lets those models decide.
PROOF illustrationmeasured results sentiment task · nine time windows · six honest, three malicious · Sci Rep 15:9177
How often the reward was right same nine windows as the animation
0%50% · CHANCE100%
One honest modela single clean time window
90.26%
Everything merged into oneall nine windows pooled · no partition
48.78%
COBRAnine models · combined by declared trust
72.45%
COBRA recovers most of the collapse, not all of it — the untrusted models are still in the vote.
ALSO REPORTEDBest configuration reaches 85% (AVGA, twenty windows) — a different setup from the three bars above.
SECOND TASKOpen-ended conversation, same method: ≈30% improvement over the unprotected reward.
SCALEReward models tested up to GPT-2 XL · 1.5B parameters.
people rate answers keep the time periods apart one model each weight by trust a reward you can use
Preserve the signal.
Suppress the corruption.
Malicious feedback should not become the reward.
HAIDER · RAHMAN · DEVABHAKTUNI · MOEYKENS · CHAKRABORTY UNIVERSITY OF MAINE · ILLINOIS STATE Sci Rep 15, 9177 (2025) ↗
Paper ↗
0:00 / 0:50
NOW TEACHING