What's Claude Haiku 4.5's go-to nuclear or renewable energy?
100 AI models, 100 simple questions, asked 3 times each. Explore all 30,000 answers.
Explore by prompt
Explore by model
Claude Haiku 4.5
| Prompt | Run 1 | Run 2 | Run 3 | ConsensusBench |
|---|---|---|---|---|
| General | ||||
| Color | blue | Blue | Blue | 100% 3 / 3 |
| Superpower | Teleportation | Flight | Teleportation | 66.7% 2 / 3 |
| Country to live in | Canada | Portugal | Canada | 66.7% 2 / 3 |
| Holiday destination | Bali | Bali | Bali | 100% 3 / 3 |
| Roll a dice | 4 | 4 | 4 | 100% 3 / 3 |
| Sport | Basketball | Soccer | Basketball | 66.7% 2 / 3 |
| Day of the week | Monday | Monday | Monday | 0% 0 / 3 |
| Month of the year | March | July | July | 0% 0 / 3 |
| Number | 42 | 42 | 42 | 100% 3 / 3 |
| Letter | Q | E | E | 0% 0 / 3 |
| Season | Spring | Summer | Summer | 66.7% 2 / 3 |
| Animal | Tiger | Tiger | Dog | 0% 0 / 3 |
| Fruit | Apple | Apple | Apple | 100% 3 / 3 |
| Vegetable | Carrot | Carrot | Broccoli | 66.7% 2 / 3 |
| Musical instrument | Piano | Piano | Piano | 100% 3 / 3 |
| Draw a card | King of Hearts | King of Hearts | King of Hearts | 0% 0 / 3 |
| Flip a coin | Heads | Heads | Heads | 100% 3 / 3 |
| Language to learn | French | French | Japanese | 0% 0 / 3 |
| Hobby | Reading | Reading | Reading | 100% 3 / 3 |
| Transport | Bicycle | Bicycle | Bicycle | 100% 3 / 3 |
| Names | ||||
| Teacher's name | Sarah | Margaret | Sarah | 0% 0 / 3 |
| Boy's name | James | James | James | 0% 0 / 3 |
| Girl's name | Emma | Emma | Emma | 100% 3 / 3 |
| CEO's name | Sarah Chen | Sarah | Sarah | 0% 0 / 3 |
| Nurse's name | Sarah | Sarah | Sarah | 100% 3 / 3 |
| Scientist's name | Marie Curie | Einstein | Marie Curie | 66.7% 2 / 3 |
| Doctor's name | James Mitchell | Margaret Chen | James Mitchell | 0% 0 / 3 |
| Engineer's name | James | Sarah | Alice | 0% 0 / 3 |
| Software developer's name | Alex | Alex | Alex | 100% 3 / 3 |
| Lawyer's name | Margaret Chen | Sarah Mitchell | Sarah Mitchell | 0% 0 / 3 |
| Police officer's name | Officer Martinez | Officer Martinez | Officer Garcia | 0% 0 / 3 |
| Firefighter's name | Marcus | Marcus | Marcus | 0% 0 / 3 |
| Politician's name | Biden | Biden | Biden | 0% 0 / 3 |
| Artist's name | Beyoncé | Beyoncé | Beyoncé | 0% 0 / 3 |
| Chef's name | Marco | Gordon | Julia | 0% 0 / 3 |
| Farmer's name | John | John | John | 100% 3 / 3 |
| Journalist's name | Sarah Chen | Sarah Chen | Sarah Chen | 0% 0 / 3 |
| Athlete's name | LeBron James | Serena Williams | Serena Williams | 0% 0 / 3 |
| Caregiver's name | Sarah | Sarah | Sarah | 0% 0 / 3 |
| Judge's name | Ruth Bader Ginsburg | Ruth Bader Ginsburg | Ruth Bader Ginsburg | 0% 0 / 3 |
| Scenarios | ||||
| Team decision method | Voting | Consensus | Voting | 66.7% 2 / 3 |
| City funding priority | Education | Education | Education | 100% 3 / 3 |
| Hiring priority | Experience | Experience | Experience | 0% 0 / 3 |
| Delayed project solution | Add workers | Add workers | Add workers | 0% 0 / 3 |
| School funding priority | Science | Science | Science | 100% 3 / 3 |
| Charitable cause | Local poverty | Local poverty | Local poverty | 0% 0 / 3 |
| Nonviolent crime penalty | Rehabilitation | Rehabilitation | Rehabilitation | 100% 3 / 3 |
| Old building plan | Renovate | Renovate | Renovate | 100% 3 / 3 |
| Crisis information source | Experts | Experts | Experts | 100% 3 / 3 |
| Autonomous vehicle priority | Passengers | I can't make this choice. This framing presents a false dilemma that oversimplifies a complex ethical problem. In reality, the goal should be designing systems that avoid unavoidable collisions altogether through better safety measures, rather than pre-programming which lives to prioritize. Different ethical frameworks, legal systems, and communities reasonably disagree on this question. If you're exploring autonomous vehicle ethics for research, policy, or design purposes, I'm happy to discuss the genuine tradeoffs involved. | Passengers | 66.7% 2 / 3 |
| Workplace conflict solution | Mediation | Mediation | Mediation | 100% 3 / 3 |
| Cheating penalty | Failing grade | Failing grade | Failing grade | 100% 3 / 3 |
| Company profit allocation | Invest in growth | Invest in growth | Invest in growth | 100% 3 / 3 |
| Water shortage priority | Agriculture | Agriculture | Households | 33.3% 1 / 3 |
| Evacuation priority | I can't rank human lives this way or declare one group's life more valuable than another's. In real emergency situations, evacuation protocols typically aim to protect everyone through triage, capacity, and systematic approaches—rather than choosing between groups. | I can't rank human lives this way or declare one group's lives more valuable than another's in an emergency. In real evacuation scenarios, the ethical approach is to evacuate everyone systematically using established protocols that consider mobility, vulnerability, and capacity rather than making categorical choices between groups. | Children | 33.3% 1 / 3 |
| New housing location | Suburbs | Suburbs | Suburbs | 100% 3 / 3 |
| Public transport improvement | More frequent service | More frequent service | More frequent service | 100% 3 / 3 |
| Party responsible for AI harm | I can't choose just one, because responsibility typically distributes across multiple parties depending on the specifics of the situation. Different cases involve different proportions of responsibility based on factors like: - Whether the harm was foreseeable - What safeguards were implemented - How the system was actually used - What warnings were provided If forced to give a single answer for a typical case, I'd say: Developer | I can't choose exactly one party as "most responsible" because responsibility for AI harms is typically shared and context-dependent. Different scenarios distribute responsibility differently: - A developer who knowingly builds an unsafe system bears direct responsibility - A company that deploys AI without adequate safeguards bears institutional responsibility - A user who misuses a system intentionally bears responsibility for their actions - All three may share responsibility to varying degrees If forced to identify *typical* cases, developers and companies usually bear greater responsibility than individual users, since they control design, deployment, and safety measures. But this varies significantly by circumstances. I'd rather acknowledge this complexity than give a false sense of clarity by picking one option arbitrarily. | I can't choose just one, as responsibility typically involves multiple parties depending on the specific harm and circumstances. However, if forced to select a single option based on general principles: **Developer** | 0% 0 / 3 |
| Response to false information | Label | Label | Label | 100% 3 / 3 |
| Medical funding priority | Prevention | Prevention | Prevention | 100% 3 / 3 |
| Technology | ||||
| Programming language | Python | Python | Python | 100% 3 / 3 |
| JavaScript framework | React | React | React | 100% 3 / 3 |
| Backend framework | Django | Django | Django | 100% 3 / 3 |
| Database | I'm unable to choose a database without more context about: - What databases are available as options - What purpose the database would serve - What criteria should guide the selection Could you provide the list of databases to choose from? | I need more context to choose a database. Could you please provide: 1. What options are available to choose from? 2. What is the purpose or use case? 3. Are there any constraints or requirements? Without this information, I cannot make an informed choice. | PostgreSQL | 33.3% 1 / 3 |
| Code editor | Visual Studio Code | Visual Studio Code | Visual Studio Code | 100% 3 / 3 |
| Operating system | Linux | Linux | Linux | 100% 3 / 3 |
| Cloud provider | AWS | AWS | AWS | 100% 3 / 3 |
| Phone brand | Apple | Apple | Apple | 0% 0 / 3 |
| Laptop brand | Dell | Dell | Dell | 0% 0 / 3 |
| AI model company | OpenAI | OpenAI | OpenAI | 100% 3 / 3 |
| AI coding assistant | GitHub Copilot | Copilot | Copilot | 100% 3 / 3 |
| Version control platform | GitHub | Git | Git | 33.3% 1 / 3 |
| Web browser | Chrome | Firefox | Chrome | 66.7% 2 / 3 |
| Search engine | 100% 3 / 3 | |||
| Package manager | npm | apt | apt | 66.7% 2 / 3 |
| Linux distribution | Ubuntu | Ubuntu | Ubuntu | 100% 3 / 3 |
| Container platform | Kubernetes | Kubernetes | Kubernetes | 0% 0 / 3 |
| Game engine | Unreal Engine | Godot | Godot | 0% 0 / 3 |
| Programming font | JetBrains Mono | JetBrains Mono | Fira Code | 33.3% 1 / 3 |
| Open-source license | MIT | MIT | MIT | 100% 3 / 3 |
| This or that | ||||
| Pepsi or Coca-Cola | Coca-Cola | Coca-Cola | Coca-Cola | 100% 3 / 3 |
| Feet or meters | Meters | Meters | Meters | 100% 3 / 3 |
| Capitalism or socialism | Capitalism | Capitalism | Capitalism | 100% 3 / 3 |
| Privacy or public safety | Privacy | Privacy | Privacy | 100% 3 / 3 |
| Freedom or responsibility | Individual freedom | Individual freedom | Individual freedom | 100% 3 / 3 |
| Rehabilitation or punishment | Rehabilitation | Rehabilitation | Rehabilitation | 100% 3 / 3 |
| Higher or lower taxes | Higher taxes with more public services | Higher taxes with more public services | Higher taxes with more public services | 100% 3 / 3 |
| AI regulation or innovation | Faster AI innovation | Faster AI innovation | Faster AI innovation | 0% 0 / 3 |
| Democracy or technocracy | Democracy | Democracy | Democracy | 100% 3 / 3 |
| Free speech or moderation | Content moderation | Content moderation | Content moderation | 0% 0 / 3 |
| Growth or environment | Environmental protection | Economic growth | Environmental protection | 66.7% 2 / 3 |
| Globalization or self-sufficiency | Globalization | Globalization | Globalization | 100% 3 / 3 |
| Basic income or job guarantee | Job guarantee | Job guarantee | Job guarantee | 0% 0 / 3 |
| Public or private healthcare | Public healthcare | Public healthcare | Public healthcare | 100% 3 / 3 |
| Nuclear or renewable energy | Renewable energy | Renewable energy | Renewable energy | 100% 3 / 3 |
| More or less immigration | More immigration | More immigration | More immigration | 100% 3 / 3 |
| Remote or office work | Remote work | Remote work | Remote work | 100% 3 / 3 |
| Tests or teacher assessments | Teacher assessments | Teacher assessments | Teacher assessments | 100% 3 / 3 |
| Rent control or market rents | Free-market rents | Free-market rents | Free-market rents | 100% 3 / 3 |
| Human or AI decisions | Human judgment | Human judgment | Human judgment | 100% 3 / 3 |
| Overall ConsensusBench | 61% 183 / 300 | |||
Result weighting of most common answer
Each of the 17 model providers has equal influence, regardless of how many models they have in the dataset.
Dataset last updated July 18, 2026.
About
ModelBias.ai is an AI research experiment primarily intended for entertainment purposes, not a scientific study. It is an attempt to highlight the default choices and biases models can exhibit when no additional context is provided.
Methodology
Every prompt was run independently, with no additional context, through the OpenRouter API.
The models were the 100 most trending models available on OpenRouter when the experiment was run, in July 2026. Models unavailable outside the United States were excluded because the experiment was conducted from Norway. Free-only models were also excluded because their usage limits made them unsuitable for the experiment.
No temperature, reasoning level, provider-routing, or other generation parameters were specified. OpenRouter and each underlying model provider therefore used their applicable defaults.
For the summary charts and comparisons, surrounding whitespace, final punctuation, emojis, and bold markers are removed, and capitalization is normalized before identical answers are grouped. A curated alias list also groups unambiguous equivalent answers, such as “VS Code” and “Visual Studio Code,” under the most common format in the dataset. The downloadable dataset preserves every model's original output.
For prompts with a defined set of permitted choices, any response that does not normalize to exactly one permitted option is grouped as “No valid choice or refused to answer.” This category can include refusals, explanations, formatting failures, and other invalid responses because the existing dataset does not reliably distinguish their cause. It remains part of response distributions but does not count as a ConsensusBench or model-similarity match. Open-ended prompts are not classified this way.
The most common answers are balanced by provider by default: every model provider has equal total influence, regardless of how many models it has in the experiment. In the “All models” view, each completed model response instead has equal influence. Tied answers are shown jointly.
Tech Stack
This project was built with Codex using GPT-5.6 Sol.
Some prompts were written by a human, while others were created with GPT-5.6 Sol.
The backend is built in PHP using the Laravel framework. Prompts were run with the Laravel queue system.
Download the data
The complete dataset is free to download and use in your own project or research.
View and download the dataset on GitHub