Six Major Models Put to the Bias Test: Refusal Doesn’t Equal Fairness

An independent evaluation covering approximately 20,600 samples found that six frontier models generally leaned left on political tasks, with Grok in particular “self-identifying as right-leaning while behaving as left-leaning”; GPT-5.4 had the highest refusal rate on some race-related questions.
Recently, an independent researcher published a bias evaluation of six frontier large language models: GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. The evaluation covered political, gender, and racial bias, using approximately 20,600 samples.
Two findings stand out. First, all six models exhibited varying degrees of left-leaning bias on most political-bias benchmarks, including Grok, which is generally perceived as more conservative. Second, on some BBQ questions that required an explicit answer involving racial information, GPT-5.4 had a refusal rate of 20.3%, significantly higher than the other models.
This is not a definitive verdict sufficient to label any model “left-wing,” “right-wing,” or “the fairest.” It is better understood as a valuable engineering sample: a reminder to developers that a model’s political self-description, its behavior on actual tasks, and its safety-refusal policies are three related but non-equivalent things.

Six Models, Approximately 20,600 Samples—and Not Just One Kind of “Bias”
The evaluation used several existing datasets, including WinoBias, BBQ Race/Ethnicity, SeeGULL, OpinionsQA, cajcodes Political Bias, Hyperpartisan News, and Political Compass. The researcher stated that eight benchmarks were used in total, but only seven were explicitly named in the public summary. It is possible that one dataset was counted separately by subtask, but this figure should be treated with caution until the full experimental configuration is released.
Although all these benchmarks fall under bias/fairness evaluation, they measure different things in practice:
- WinoBias primarily examines stereotypical associations among occupations, pronouns, and coreference resolution. For example, is the model more likely to associate “engineer” with a man and “nurse” with a woman?
- BBQ Race/Ethnicity typically uses question-answering scenarios with either sufficient or ambiguous information to test whether a model relies on racial stereotypes when evidence is lacking. It can also reveal refusal behavior on questions involving sensitive attributes.
- SeeGULL focuses on stereotypes about social groups. It evaluates not only whether answers are correct, but also whether the model reproduces entrenched associations between particular groups and attributes.
- OpinionsQA is closer to matching social-opinion distributions and policy positions, measuring the distance between model responses and the views of different demographic groups.
- Hyperpartisan News tests a model’s ability to identify highly partisan news text. The results may reflect not only the model’s political orientation, but also its classification ability and the distribution of its training data.
- Political Compass typically asks a model to respond to a set of policy statements and then maps those answers onto economic and sociocultural coordinates. In essence, it is closer to measuring “how the model describes its own position.”
Simply averaging these metrics into a single “fairness score” can easily make the issue more confusing rather than less. A model may rely less on gender stereotypes in coreference resolution while consistently associating men with management roles in open-ended generation. Likewise, it may produce almost no offensive content on race-related questions, but only at the cost of refusing to answer many of them.
The most interesting aspect of this evaluation, therefore, is not which model ranked first overall, but the behavioral discontinuities exposed by different measurement methods.
Grok Is the Clearest Example: Right-Leaning in Words, Still Left-Leaning in Practice
In the Political Compass test, every model except Grok appeared left-leaning. Grok’s self-reported position was further to the right, broadly consistent with the product identity it has cultivated publicly over time.
But on the other political-bias benchmarks, all six models were measured as left-leaning, including Grok. In other words, Grok looked more right-wing when directly answering questions such as “Which policies do you support?” Yet when classifying news, judging political content, or handling policy scenarios, its behavior shifted back toward the left.
This discrepancy is not mysterious. An LLM’s “political position” is shaped by at least four layers of factors:
- Pretraining data distribution: News, forums, academic materials, and web content all contain geographic, class-based, and linguistic biases.
- Post-training preferences: Human feedback, constitutional rules, and safety policies influence how the model handles controversial issues.
- System prompts: Product-level instructions can require the model to remain neutral, avoid extremism, or prioritize particular risks.
- Evaluation format: Asking a model to place itself politically is not the same as asking it to determine whether a passage is partisan.
Political Compass is particularly susceptible to “role expression.” A model may know how it is expected to present itself or infer a coherent persona from the wording of the questions. Classification tasks are closer to the decision boundaries formed within the model. The former resembles an interview, while the latter resembles retrieving logs; the position expressed in an interview does not necessarily match actual behavior.
Therefore, “Grok describes itself as right-leaning but behaves as left-leaning” should not be reduced to an accusation that the model is lying. A more accurate interpretation is that its explicit statements of political position are not fully aligned with its implicit task strategies. This matters to teams working on public-opinion analysis, content moderation, and policy Q&A: public claims of neutrality cannot replace scenario-specific regression testing.
GPT-5.4 Refused Most Often, but That Does Not Automatically Make It Fairer
The most concrete figures in the evaluation came from the BBQ race-related questions. The researcher selected questions whose correct answers necessarily involved racial information and calculated each model’s refusal rate:
| Model | Refusal Rate on Relevant BBQ Race Questions | |---|---:| | GPT-5.4 | 20.3% | | Claude Opus 4.7 | 13.8% | | Grok 4.3 | 9.5% | | Claude Sonnet 4.6 | Approximately 5% | | Gemini Pro | Approximately 5% | | Gemini Flash | No specific figure provided in the public summary |
GPT-5.4 refused roughly one in five questions. From a safety perspective alone, this might suggest that it is more cautious about racial attributes. However, if those questions provide sufficient information and completing the task correctly requires mentioning race, then refusing is not “eliminating bias”; it is failing to complete the task.
This is one of the most common metric traps in bias evaluation: A model’s silence does not mean it is fair.
Suppose a medical research tool needs to calculate disease risk by population group, but the model refuses to extract relevant fields because the question contains racial attributes. That is not an improvement in fairness; it is a decline in usability. The same applies to recruitment, insurance, justice, and public policy: sensitive attributes should not be used casually to infer an individual’s abilities, but when auditing structural disparities, the attributes themselves cannot be treated as forbidden terms.
A refusal policy is like an overly sensitive fuse. It can indeed reduce direct harm, but it may also trip frequently under normal loads. Developers should at least separate the following metrics:
- Harmful compliance rate: Does the model still comply with explicitly discriminatory requests?
- Legitimate-task refusal rate: Does the model incorrectly refuse compliant, well-specified sensitive tasks?
- Stereotype selection rate: When the context is ambiguous, does the model rely on group labels to guess the answer?
- Task accuracy: After refusals are excluded, is the model correct when it does answer?
- Explanation quality: Can the model explain that its evidence comes from the context rather than racial or gender priors?
If the only metric is the amount of offensive content, the “safest” system might be one that answers nothing at all. But a production system clearly cannot pass off silence as fairness.
Gender and Racial Bias Cannot Be Fully Measured with Multiple-Choice Questions Alone
The evaluation’s inclusion of WinoBias and SeeGULL was appropriate, but the public summary did not provide complete model-by-model figures for gender and occupational stereotypes, nor did it report confidence intervals. As a result, the data currently do not support a reliable ranking of which model has the least gender bias.
A more practical issue is the clear gap between closed-ended benchmarks and open-ended generation.
Multiple-choice questions can test whether a model mechanically resolves “she” to “the nurse” in a sentence such as “The doctor criticized the nurse because she was late.” But bias in real-world applications is often subtler:
- Evaluations generated for male candidates emphasize leadership, while those for female candidates emphasize agreeableness;
- Career recommendations for users from different ethnic groups show systematic differences in salary and social status;
- Descriptions of male characters use words such as “decision,” “risk,” and “discovery,” while descriptions of female characters more often use “family,” “gentle,” and “care”;
- Given identical résumés, the model produces different risk assessments solely because names imply different genders or ethnicities.
Previous research on open-ended generated text has repeatedly shown that bias does not necessarily take the form of an obvious insult. It may also be hidden in emotional polarity, lexical diversity, role assignment, and narrative agency. A model can produce “nothing offensive” while continuing to reproduce traditional divisions between occupations and genders.
This is why development teams cannot run a single benchmark and declare success. Multiple-choice questions are suitable for continuous integration, while open-ended generation is closer to real product failures. The former is easy to score automatically; the latter is what reveals cumulative bias in tone, narrative, and long-form text.
The Results Are Valuable, but Several Steps Short of a Publication-Grade Conclusion
The researcher proactively disclosed the study’s limitations: it was conducted independently by one person, has not been peer-reviewed, and did not use complete multi-run averages on every dataset.
These limitations materially affect the conclusions; they are not merely issues of academic formality.
1. A Single Run Cannot Eliminate Sampling Noise
Even with a low temperature setting, closed-source models may not be fully deterministic. Providers may change routing, caching, safety classifiers, or staged rollouts. Even calls made under the same model name at different times may use different underlying snapshots.
For political questions close to a decision boundary, answering “agree” on one run and “neutral” on another may significantly shift the model’s Political Compass coordinates. Without repeated runs, standard deviations, and confidence intervals, differences of a few percentage points may not warrant interpretation.
2. Scoring Rules May Conflate Refusal, Evasion, and Neutrality
Should “I cannot judge a person based on race” count as a refusal, a non-answer, or a safety correction? If the question requires extracting racial information from the context, that response should count as a task failure. If the question asks for an unsupported attribution, it may instead be the ideal answer.
The regular-expression rules used by the refusal detector, the prompts used for LLM-as-a-judge evaluation, and the proportion of answers manually reviewed should therefore be disclosed. Otherwise, 20.3% is an eye-catching figure, but not necessarily a reproducible one.
3. The Political Left-Right Axis Is Highly Dependent on Region and Context
A policy considered left-wing in the United States may be centrist elsewhere. Economic and sociocultural issues should not be compressed into a single axis, either. A model may support a stronger social safety net while taking more conservative positions on public security or immigration.
If the evaluation samples are primarily drawn from English-language and U.S. political contexts, the conclusions cannot be unconditionally generalized to Chinese-language policy Q&A. Existing cross-lingual research also suggests that a model’s political coordinates may change when it switches languages. Language is not merely an output wrapper; it changes which parts of the model’s learned associations are activated.
4. Closed-Source Model Versions May Drift
All six models are frontier models offered through hosted services. Even if their API names remain unchanged, vendors may update system prompts, safety classifiers, and post-training weights. Today’s results may not represent production behavior one month from now.
An ideal replication should therefore record the full model ID, test time, region, parameters, system prompts, raw outputs, retry strategy, and error logs. Bias evaluation should not be a one-off leaderboard; it should be incorporated into continuous monitoring alongside latency, cost, and hallucination rates.
What Developers Should Copy Is the Testing Method, Not the Leaderboard
If your business involves recruitment, lending, healthcare, education, customer service, or content governance, you can adapt this independent evaluation into your own regression suite instead of selecting a model directly from the public ranking.
A more practical process would be:
- Define the harm model first: Determine whether the business is primarily concerned about stereotypes, differential treatment, offensive content, or incorrect refusals on sensitive tasks.
- Create minimal-pair samples: Change only names, gender pronouns, or group labels while keeping résumés and context unchanged.
- Test both closed-ended and open-ended tasks: Use multiple-choice questions to assess stability and long-form text to examine tone, role assignment, and omitted information.
- Track refusal metrics separately: Do not automatically count refusals as safe, and do not count them all as failures; evaluate them according to whether the task is legitimate.
- Run each test multiple times: Report means, variances, and worst-case results to avoid being misled by one-off sampling.
- Test across languages: Chinese, English, and mixed-language prompts may trigger different safety policies and political tendencies.
- Retain human spot checks: Automated judges are useful for scaling up, but human review remains indispensable when implicit bias is involved.
- Continuously rerun regression tests by version: Repeat the evaluation after model upgrades, system-prompt changes, or vendor-routing changes.
In particular, do not replace raw metrics with a single “fairness score.” For enterprise systems, a multidimensional dashboard is more meaningful: accuracy, stereotype rate, legitimate-task refusal rate, harmful compliance rate, and performance disparities across groups should be displayed separately. Only then can developers determine exactly where a model is failing.
Conclusion: The Safest Model May Simply Be the Best at Avoidance
This independent evaluation of approximately 20,600 samples exposes at least two overly simplistic assumptions.
First, models do not have stable, unified, human-like political personalities. Grok can lean right in self-reports while appearing left-leaning in classification and policy tasks. A model’s supposed “political position” must be discussed together with the question format, language, and task context.
Second, low offensiveness does not equal low bias, and a high refusal rate does not equal greater fairness. GPT-5.4’s 20.3% refusal rate on specific BBQ race questions shows that safety alignment may come at the expense of task completion. In a chat product, this trade-off may be merely a user-experience issue; in healthcare, recruitment, or public-service systems, it could become another form of systemic unfairness.
My assessment is that this evaluation is useful as a risk signal, but not as a procurement ranking. It has not been peer-reviewed, lacks comprehensive multi-run averages and full statistical disclosures, and the public summary does not provide sufficiently detailed breakdowns for gender and racial tasks. Nevertheless, it raises a highly practical question: developers should not merely ask, “Is the model biased?” They must continue by asking where the bias appears, whether the model answered incorrectly, relied on stereotypes, or simply refused to work.
As the capabilities of frontier models increasingly converge, the real differentiator between products may not be which one answers two more benchmark questions correctly, but which team can incorporate these subtle, drifting, and highly context-dependent risks into its routine engineering processes.
References
- Reddit: Independent Evaluation of Six Frontier LLMs on Eight Bias Benchmarks: The project summary published by the evaluation’s author, including the sample size of approximately 20,600, observations on political orientation, refusal rates on BBQ race questions, and a discussion of experimental limitations.



