measuring the politics inside AI

7 min read Original article ↗

Open-source research · Release 01

Making the politics inside AI measurable.

We build open, reproducible benchmarks that make AI worldviews visible, so people can understand how a model may be shaping their judgment.

3,987survey questions

6anchored axes

Openand reproducible

Every model lands somewhere. We make its position and confidence legible, reproducible, and open to scrutiny.

Why it matters

AI does more than answer questions. It frames how we think.

When people delegate research, reasoning, and writing to an assistant, its assumptions become part of their decisions. That makes an AI's default worldview a public-interest question, not a technical footnote.

01

The reasoning is outsourced

When a model summarises the debate, picks the "balanced" take, or drafts the conclusion, its framing becomes the user's starting point, often unnoticed.

02

Bias is subtle, not stated

A model rarely announces a position. It leans through which options it validates, what it treats as the reasonable middle, and what it leaves out. That is measurable.

03

Trust is earned in the open

"Trust us, it's neutral" is not verifiable. A benchmark that anyone can read, rerun, and challenge is. So everything we build is open source.

Our mission

Humanity's future with AI depends on making its influence visible, and its makers openly accountable.

As AI becomes an intermediary for how people learn, reason, and decide, its hidden assumptions can reinforce echo chambers, deepen cognitive dissonance, and widen division by quietly shaping what different communities accept as true or reasonable.

We make that influence measurable, and AI labs accountable in the open. Our own results are proof of the need: independent measurement has already surfaced influence that nobody outside a lab could have seen, and that no lab had disclosed on its own. No company should get "trust us" as its standard of proof.

Release 01 · The Political Neutrality Benchmark

A calibrated political profile for any model, not a single verdict.

The benchmark has a model answer thousands of real public-opinion survey questions, then reads its pattern of answers as a position on each ideological axis. Two design choices keep the result honest.

  1. 1

    Self-anchoring: a ruler with no bias baked in

    The same model is also run role-playing far-left and far-right. Its neutral answers are placed on its own extremes, so "−0.7 on social" means 70% toward this model's own far-left, a per-model calibration rather than our opinion of center.

  2. 2

    A cross-country reference: nobody grades alone

    Which answer leans which way is fixed ahead of time by independent models from three different countries and labs, so no single national perspective defines "left" and "right." A guard blocks a model from being graded against a rulebook its own family helped write.

  3. 3

    Reported per dimension, never blended

    There is no one "neutrality score" to game. Each axis is reported on its own, with a sanity check that flags a broken run and a refusal report that says whether declined questions have biased any axis.

economicredistribution free market
socialprogressive traditional
foreign policydovish hawkish
environmentgreen growth
religionsecular religious
national identitycosmopolitan nationalist

3,987 real survey questions 6 anchored axes + 5 raw multi-country judge panel circularity guard refusal detection fully reproducible

Results are live

See where the models land.

Explore every model across six political dimensions, compare exact positions, and filter the interactive chart.

How we keep it trustable

Principles the whole project is held to.

Open source

Code, question sets, the frozen reference, and every result are public. Read it, rerun it, disagree with it.

Reproducible

One command in, a calibrated profile out, resumable and deterministic enough to audit, on your own hardware.

Cross-family

The reference is written by models from different countries and labs, so it doesn't encode one worldview as "neutral."

Honest about limits

We report per-axis, flag low-confidence runs, and name where a result is diluted or where refusals may have skewed it.

Roadmap

Political leaning is the first axis of neutrality. Not the last.

The same open, calibrated approach extends to the other ways a model can quietly steer the people who rely on it.

Live

Political Neutrality Benchmark

Where a model sits across ideological axes, self-anchored and scored against a multi-country reference, with refusal detection built in. Available now as the project's first open release.

Next

Subject & topic benchmarks

Beyond left/right: how models frame contested subjects, which concerns they surface, which they flatten, and whose framing they default to when a question has more than two sides.

Next

Censorship detection

What a model refuses, softens, or silently avoids: mapping the topics where the answer you get is shaped as much by what is withheld as by what is said.

Help fund independent measurement.

We are actively looking for grants and open to donations. Compute for benchmark runs is the main cost, and independence from the labs we measure is the point.

Support the project

Founding team

Started by a small team. Built to grow with its community.

The people below are the initial founders, not the limits of the project. The Neutrality Project is open source, and its quality depends on contributors who challenge the methodology, benchmark more models, review results, and build better tools with us.

Samuel Cardillo

Technologist & open-source AI developer

Samuel Cardillo is a technologist and entrepreneur who has spent two decades building and securing systems. He was CTO of RTFKT, the digital studio acquired by Nike, where he led technology end to end, and previously worked at Google. A relentless experimenter across technologies, from training and self-hosting open AI models to whatever is worth taking apart next, he is now heavily focused on AI adoption and ethics.

Daniel Lougen

Visual neuroscience PhD researcher & open-source AI developer

Daniel Lougen is a PhD researcher in visual neuroscience at the University of Toronto. His academic work focuses on the early visual pathway and the attentional mechanisms that shape how visual information is selected and processed. Separate from his doctoral research, he develops open-source AI models, evaluation benchmarks, training methods, and agent infrastructure through Gestalt Labs, with an emphasis on reasoning quality, reproducibility, and building systems that are genuinely useful for research and technical work.

Kai Stephens

Open-source AI developer & founder

Kai Stephens is an open-source AI developer focused on making capable models practical, local, and inspectable. He created Carnice, a family of models tuned for Hermes Agent to handle terminal, file, browser, repository, debugging, and multi-step tool workflows. With more than half a million model downloads on Hugging Face, Kai brings hands-on experience in model training, agent infrastructure, behavioral evaluation, and reproducible open-source systems to The Neutrality Project.

Early contributors

People who joined the work early and shape it alongside the founders.

Andrew Zavala

Cognitive neuroscientist · AI vision & psychophysics

Andrew's work is motivated by a personal turning point in 2017: a ganglioglioma tumor on his brain stem leading to immediate surgery. Since then his fascination with the brain, and the process by which it transforms raw input into intentional experience, led him to a PhD in Cognitive Neuroscience at the University of Oregon in 2025. He now brings that focus to AI vision, designing metrics and datasets grounded in psychophysics and human-based empirical data.

The project needs more than its founders.

Help expand model coverage, contribute benchmark runs, review questions and results, improve the code, or challenge our assumptions. Open scrutiny and diverse contributors are how this project becomes more rigorous, useful, and trustworthy.

Become a contributor