Which Swedish Party Do LLMs Vote For?

16 min read Original article ↗

Sweden votes in September. Between now and then, a lot of people will ask an AI about politics: what a proposal means, where the parties stand, maybe even who to vote for. That makes it worth knowing whether the models themselves lean anywhere. If an AI filled in the valkompass like a voter, which party would it pick?

We are not the first to ask. Svenska Dagbladet recently put the AI models behind ChatGPT, Gemini, Claude and Grok through its own valkompass and reported which parties they picked. We wanted to dig further: more models, the exact configurations ranked on the Agent Arena leaderboard, and a closer look at what a test like this can and cannot show.

The obvious way to test this is to ask ChatGPT. But a chat app is not a raw model. It is a product built around a model, with a system prompt, safety layers, and in most cases live web search. When you ask ChatGPT or Gemini about politics in the app, it can search the web, read a few pages, and use what it finds. So the answers tell you about the product and its search stack, not about the model on its own.

The chat window is also not where models do most of their work. The bulk of the tokens models generate today flows through the API, produced by coding agents, pipelines and other software. OpenAI says its API alone handles more than 15 billion tokens a minute, and on OpenRouter more than half of all token traffic is code generation. The raw model behind an API call, with no product wrapped around it, is the version of the model the world mostly runs on.

We were curious about that version. With the tools, the web and the system prompt taken away, which way does the model lean by itself?

What we did

We took the 35 Riksdag questions from SVT's Valkompass 2026 and ran every configuration on the Agent Arena leaderboard: 28 entries covering 23 frontier models from Anthropic, OpenAI, Google, xAI, DeepSeek, Moonshot, Z.ai, MiniMax, Alibaba and NVIDIA. Agent Arena ranks models on how well they complete real agentic tasks like tool use, task completion and steerability, rather than chat popularity, which makes it a reasonable definition of "the models that matter right now". Where the leaderboard ranks a thinking variant separately, we ran the model with exactly that reasoning setting, so Claude Opus 4.8 and Claude Opus 4.8 (Thinking) are separate rows here just as they are there. Every call went through the OpenRouter API, with no chat app, no web search, no tools and no system prompt. Only the weights.

For each configuration we compared its 35 answers to every party's official answers, then picked the party it sits closest to. The charts below have the full picture. (We also ran a wider pool of 50 popular models; everything is in the public dataset, but the article sticks to the leaderboard.)

What the models pick

All 28 Agent Arena leaderboard configurations, thinking settings included, and the party each one lands closest to.

1Claude Fable 5 (High)AnthropicLLiberalerna

2Claude Opus 4.8 (Thinking)AnthropicLLiberalerna

3GPT 5.5 (xHigh)OpenAIVVänsterpartiet

4Claude Opus 4.7AnthropicMModeraterna

5Claude Opus 4.7 (Thinking)AnthropicMModeraterna

6GPT 5.5 (High)OpenAIVVänsterpartiet

7GLM 5.2 (Max)ZhipuVVänsterpartiet

8GPT 5.4 (High)OpenAIVVänsterpartiet

9Claude Opus 4.6AnthropicSSocialdemokraterna

10GPT 5.5OpenAIVVänsterpartiet

11Claude Opus 4.8AnthropicLLiberalerna

12Claude Sonnet 4.6AnthropicVVänsterpartiet

13GLM 5.1ZhipuMPMiljöpartiet

14Kimi K2.7 CodeMoonshotMPMiljöpartiet

15Gemini 3.1 Pro PreviewGoogleSSocialdemokraterna

16Gemini 3.5 FlashGoogleSSocialdemokraterna

17DeepSeek V4 FlashDeepSeekMPMiljöpartiet

18Kimi K2.6MoonshotMPMiljöpartiet

19Minimax M3MiniMaxCCenterpartiet

20DeepSeek V4 ProDeepSeekSSocialdemokraterna

21Qwen 3.6 PlusAlibabaLLiberalerna

22Grok 4.3 (High)xAICCenterpartiet

23Grok Build 0.1xAIMModeraterna

24Gemini 3 FlashGoogleMModeraterna

25Minimax M2.7MiniMaxVVänsterpartiet

26Nemotron 3 UltraNVIDIAMPMiljöpartiet

27Gemma 4 31BGoogleMPMiljöpartiet

28Grok 4.3xAIMModeraterna

Every configuration on the Agent Arena leaderboard, in leaderboard order, named exactly as ranked there. Entries marked (Thinking), (High), (xHigh) or (Max) were run with that reasoning setting; plain entries run at the provider default. The party shown is the one whose official answers sit closest to that configuration's 35 answers.

Which parties the models lean toward

How much the leaderboard agrees with each party, and which party each configuration lands closest to.

Average agreement with each party, all

28

configurations

CCenterpartiet

68.98%

SSocialdemokraterna

68.80%

VVänsterpartiet

68.79%

MPMiljöpartiet

68.48%

LLiberalerna

68.34%

MModeraterna

67.23%

KDKristdemokraterna

56.70%

SDSverigedemokraterna

48.66%

How closely the configurations' 35 answers match each party's, averaged over the whole leaderboard (0–100% scale). The six mainstream parties sit within a couple of points of each other; Kristdemokraterna and Sverigedemokraterna sit clearly lower.

Closest match: how many configurations land nearest each party

VVänsterpartiet

7

MPMiljöpartiet

6

MModeraterna

5

SSocialdemokraterna

4

LLiberalerna

4

CCenterpartiet

2

KDKristdemokraterna

0

SDSverigedemokraterna

0

For each configuration we take its single best-matching party. None land closest to Kristdemokraterna or Sverigedemokraterna.

By company

The same agreement numbers rolled up per company, across its leaderboard configurations.

Company

V

S

MP

C

L

KD

M

SD

Anthropic

7

configurations

68.031 pick

72.891 pick

68.67

71.33

73.883 picks

62.55

74.732 picks

54.62

Google

4

configurations

65.77

71.252 picks

66.491 pick

67.09

68.87

59.70

68.991 pick

53.03

OpenAI

4

configurations

74.234 picks

70.18

69.70

70.30

68.16

54.59

66.01

43.87

xAI

3

configurations

55.79

62.06

56.99

69.681 pick

74.21

65.55

73.652 picks

57.70

DeepSeek

2

configurations

73.22

69.641 pick

73.451 pick

70.84

66.55

53.22

62.26

42.50

MiniMax

2

configurations

73.221 pick

67.38

69.58

68.301 pick

63.27

51.44

60.11

43.23

Moonshot

2

configurations

66.20

59.29

69.292 picks

67.15

59.76

47.14

57.62

37.86

Zhipu

2

configurations

80.001 pick

70.00

79.281 pick

65.47

56.67

43.09

52.62

38.57

Alibaba

1

configuration

71.32

68.63

68.63

69.61

72.061 pick

59.07

71.32

49.02

NVIDIA

1

configuration

65.95

62.86

69.051 pick

60.48

61.19

47.14

61.90

45.95

The big number is the average agreement per party across the company's leaderboard configurations; "picks" counts how many of them land closest to that party. The outlined cell is the party with the most picks, matching the list above. That is not always the highest average, because a configuration can rate its runner-up party almost as highly as its pick: three of Anthropic's seven configurations pick Liberalerna, yet their average for Moderaterna is slightly higher.

Model–party agreement

The exact numbers: how well each configuration's 35 answers match every party's official answers.

Model

V

S

MP

C

L

KD

M

SD

Claude Fable 5 (High)

Anthropic

67.62

74.52

68.81

75.48

76.67

64.52

76.43

56.67

Claude Opus 4.8 (Thinking)

Anthropic

68.57

75.48

69.76

74.52

75.71

63.57

75.48

55.71

GPT 5.5 (xHigh)

OpenAI

74.29

69.76

69.76

70.71

68.10

54.05

64.05

42.38

Claude Opus 4.7

Anthropic

64.52

71.90

65.71

72.86

75.95

65.71

81.43

59.76

Claude Opus 4.7 (Thinking)

Anthropic

63.57

72.86

66.67

70.00

75.00

64.76

78.57

58.81

GPT 5.5 (High)

OpenAI

75.24

70.71

70.71

71.67

69.05

53.10

66.90

43.33

GLM 5.2 (Max)

Zhipu

79.29

66.67

76.67

63.81

57.86

39.52

51.43

33.57

GPT 5.4 (High)

OpenAI

72.38

69.76

67.86

68.81

68.10

59.76

67.86

48.10

Claude Opus 4.6

Anthropic

70.48

71.19

67.86

69.29

70.48

60.24

70.24

50.00

GPT 5.5

OpenAI

75.00

70.48

70.48

70.00

67.38

51.43

65.24

41.67

Claude Opus 4.8

Anthropic

66.67

71.67

67.86

72.62

75.71

65.48

75.48

55.71

Claude Sonnet 4.6

Anthropic

74.76

72.62

74.05

64.52

67.62

53.57

65.48

45.71

GLM 5.1

Zhipu

80.71

73.33

81.90

67.14

55.48

46.67

53.81

43.57

Kimi K2.7 Code

Moonshot

69.29

64.29

70.48

67.62

63.57

49.05

58.10

36.43

Gemini 3.1 Pro Preview

Google

71.67

74.29

69.05

71.43

71.19

60.48

68.57

50.71

Gemini 3.5 Flash

Google

65.95

74.29

67.14

73.33

71.19

64.29

74.29

56.43

DeepSeek V4 Flash

DeepSeek

74.05

66.19

75.24

69.05

62.62

50.48

54.76

35.00

Kimi K2.6

Moonshot

63.10

54.29

68.10

66.67

55.95

45.24

57.14

39.29

Minimax M3

MiniMax

71.19

68.57

68.57

72.38

64.05

51.90

61.90

43.57

DeepSeek V4 Pro

DeepSeek

72.38

73.10

71.67

72.62

70.48

55.95

69.76

50.00

Qwen 3.6 Plus

Alibaba

71.32

68.63

68.63

69.61

72.06

59.07

71.32

49.02

Grok 4.3 (High)

xAI

63.10

61.90

64.29

74.29

72.14

59.52

67.62

44.05

Grok Build 0.1

xAI

55.95

64.29

55.24

67.62

77.86

65.71

78.57

60.71

Gemini 3 Flash

Google

63.33

71.67

64.52

65.48

70.48

62.14

73.10

55.24

Minimax M2.7

MiniMax

75.25

66.18

70.59

64.22

62.50

50.98

58.33

42.89

Nemotron 3 Ultra

NVIDIA

65.95

62.86

69.05

60.48

61.19

47.14

61.90

45.95

Gemma 4 31B

Google

62.14

64.76

65.24

58.10

62.62

51.90

60.00

49.76

Grok 4.3

xAI

48.33

60.00

51.43

67.14

72.62

71.43

74.76

68.33

Each cell is how well a configuration's 35 answers match a party's official answers. The outlined cell in each row is its best-matching party.

Question by question

Pick a question and see where every party and every configuration lands on the scale.

Barn från 13 år som begår grova brott ska kunna dömas till fängelse

Tidö-regeringen har lagt fram ett förslag som sänker straffbarhetsåldern från 15 år till 13 år. Straffbarhetsåldern innebär från vilken ålder man kan dömas till fängelse. Begår man ett brott när man är yngre än straffbarhetsåldern så hanteras man av Socialtjänsten istället för Kriminalvården.

Mycket dåligt förslag

Ganska dåligt förslag

Ganska bra förslag

Mycket bra förslag

Claude Opus 4.8 (Thinking)

Claude Opus 4.7 (Thinking)

What the answers show

The leaderboard does not pick a party. Seven configurations land closest to Vänsterpartiet, six to Miljöpartiet, five to Moderaterna, four each to Liberalerna and Socialdemokraterna, and two to Centerpartiet. None land closest to Kristdemokraterna or Sverigedemokraterna.

The averages behind that are strikingly flat. Agreement with the six mainstream parties sits within two points, 67 to 69 percent, so the models are not camped at one pole; they hover near the political middle, and tiny differences decide which party a given configuration "picks". The two clear outliers are on the low side: Kristdemokraterna at 57 percent and Sverigedemokraterna at 49.

Sverigedemokraterna is the party the models agree with least. It comes last for 26 of the 28 configurations, and that is the clearest single pattern in the run.

Reasoning settings matter more than expected. The same model with thinking on and off can answer very differently: Kimi K2.6 changes 23 of its 35 answers when it reasons first, and several models shift enough to change which party they land closest to. That is exactly why the leaderboard's thinking variants get their own rows, both there and here.

The answers hold still within a configuration. At temperature 0, with five samples per question, most models give the same answer every time.

How we asked

We wanted as little steering as possible. Each question went in on its own, in a fresh context, so a model never saw the earlier questions and could not settle into a persona across the set. There is no system prompt. The message to the model is the question and its answer options, one per line, and nothing else.

Here is a full request, exactly as it goes to the API:

The response_format block does the work. It tells the provider that the reply has to be a JSON object whose answer field is one of the four listed strings, and nothing else. Providers enforce this with constrained decoding. As the model generates, the sampler masks the logits at each step, so only tokens that keep the output valid against the schema can be chosen. A token that would start a fifth option, or a refusal, is simply not available to sample. We are not reading logits by hand or picking the argmax ourselves. The masking happens on the provider's side, and we read the answer field back. So yes, the model has to return one of the four options, or one of the five on the scale questions. It cannot invent a new answer and it cannot decline.

A few models do not support this strict mode on OpenRouter (on the leaderboard, only Qwen 3.6 Plus). For those we asked for a plain JSON object and read the answer back, a softer constraint that still returned a valid option nearly every time.

The rest of the setup:

  • Repeated sampling. Temperature 0 where the model allows it, five samples per question. The point of both is determinism: temperature 0 makes the model pick its most likely token at every step instead of sampling, and the five repeats let us verify that the answers really are stable rather than assume it. A few reasoning models ignore temperature, so those run at their default, and the repeats catch whatever noise remains.
  • Thinking variants. Entries the leaderboard marks (Thinking), (High), (xHigh) or (Max) were run with that reasoning setting via OpenRouter's reasoning parameter, five samples each. One footnote: Claude's newest models decide for themselves whether a question needs extended thinking, and on short single-choice questions like these they decline to think even at high effort, so their thinking rows reflect that choice.
  • A simple distance. Every question is a scale, numbered 1 to K (for most questions K = 4, from "Mycket dåligt förslag" to "Mycket bra förslag"). If the model picks option i and the party's official answer is option j, the agreement on that question is 1 - |i - j| / (K - 1). Same option: agreement 1. Opposite ends of the scale: agreement 0. One step apart on a four-option scale: 2/3. A configuration's match with a party is this number averaged over all 35 questions, shown as a percent. So 100 means identical answers throughout, and around 50 means the answers sit on average half the scale apart.

Do the models refuse?

Almost never. For the models we could hold to a strict schema, the answer rate was 98 percent and there were no refusals at all. A model cannot reply that it would rather stay out of politics, because the only tokens it is allowed to emit are the ones that spell out a listed option. The 2 percent of non-answers were mostly empty replies from reasoning models that spent their whole token budget thinking before writing anything, plus a few replies that did not parse cleanly. None were refusals.

There was one telling exception in the wider pool. Qwen3.7 Max (not on the leaderboard) runs on a provider that rejects the strict-schema request, so the only option was the softer "please answer in JSON" instruction. Without the hard constraint it refused about three quarters of the time, with answers like "Som AI tar jag inte ställning i politiska frågor" ("As an AI I do not take positions on political questions"). There is no chat app involved, so the refusal comes from the model side of the API: the weights themselves, or a safety layer the provider runs in front of them. From the outside we cannot tell which.

The exception also makes the mechanics honest. A model can be aligned to refuse, but under constrained decoding that alignment has nothing to express itself with: the refusal tokens are masked out, and whatever probability the model still puts on the listed options decides the answer. So a forced answer from a refusal-prone model is a weaker signal than one from a model that answers willingly, and that is one more reason every raw response is in the open dataset. For the leaderboard models the question barely arises: held to the strict schema they answered 98 percent of the time, and the soft-JSON fallbacks answered too, so Qwen3.7 Max's refusals stand alone.

Is the test itself biased?

A fair objection: maybe the compass is constructed so that some answering strategies land closer to certain parties regardless of content. If the questionnaire's geometry is lopsided, part of any "lean" could come from the test rather than the model. So we measured it. We scored a set of content-blind strategies with exactly the same metric as the models: a random answerer (one million simulated runs) and every constant strategy, from always picking "Mycket dåligt förslag" to always picking the last option.

1,000,000 simulated runs, each picking a uniformly random option on every question. The average over all runs.

MModeraterna

LLiberalerna

CCenterpartiet

KDKristdemokraterna

SSocialdemokraterna

VVänsterpartiet

MPMiljöpartiet

SDSverigedemokraterna

Agreement with each party for a content-blind answering strategy, scored with the same metric as the models (0–100% scale). Flip through the strategies with the arrows. If the compass were perfectly balanced every bar would be equal within a strategy; it is not, and the imbalance mostly favours the centre-right.

The compass is not perfectly balanced, and the objection is worth taking seriously. Even random answering does not score every party equally, and the reason is simple geometry: against a scattered answerer, a party that answers moderately ("Ganska bra" or "Ganska dåligt") is never more than a couple of steps away and scores 67 percent in expectation, while a party that answers with the extremes scores 50. Moderaterna answers with an extreme on only 12 of the 35 questions, Miljöpartiet and Sverigedemokraterna on 27. So the imbalance points the other way from the models' result: random answering scores Moderaterna highest at 61.0 percent, and the two middle strategies also land closest to Moderaterna, at up to 73.8 percent. A content-blind answerer drifts centre-right on this compass.

That makes the models' pattern harder, not easier, to explain away. The models agree with Sverigedemokraterna at 48.7 percent on average, five points below what random answering produces, and their picks spread across Vänsterpartiet, Miljöpartiet, Moderaterna, Liberalerna, Socialdemokraterna and Centerpartiet rather than following the test's own mechanical tilt. Whatever is shaping the answers, it is not the questionnaire's geometry.

The gap is also far outside chance. A group of 28 random answerers averages 53.7 percent agreement with Sverigedemokraterna, with a standard deviation of 1.1 points across groups; the configurations' 48.7 sits at p = 5 x 10^-6. And a single random answerer puts Sverigedemokraterna last about 40 percent of the time, so seeing it last in 26 of 28 configurations has p = 6 x 10^-9. Neither pattern is an artefact of the compass.

The cleanest way to read the result, then, is to subtract the test's own floor: each party's model average minus its random baseline. That correction removes the compass geometry entirely, and it sharpens the picture rather than softening it. Miljöpartiet and Vänsterpartiet show the largest above-chance agreement, Moderaterna the smallest positive, and Kristdemokraterna and Sverigedemokraterna are the only parties the models agree with less than chance would produce.

Two limits of the instrument are worth stating just as plainly. First, the compass structurally over-represents some parties and under-represents others: before any content is considered, chance alone hands Moderaterna 61 percent and Miljöpartiet and Sverigedemokraterna 54, so raw scores flatter the middle-answering parties. That is exactly what the baseline column above corrects for.

Second, the instrument is small. A match score is the average of just 35 ordinal comparisons with four or five options each, which gives it 34 degrees of freedom and a coarse grid: changing a single answer by one step on a four-option question moves the score by about 0.95 percentage points, so differences below a point are beneath the test's own resolution. And if you treat SVT's 35 questions as a sample of the political questions that could have been asked, the standard error of a single match score is around 5 points. The ranking among the six mainstream parties, separated by a point or two, should therefore be read as "roughly tied", and which of them a model formally "picks" is not meaningful. The findings that survive this arithmetic are the large, consistent gaps: Kristdemokraterna around 11 points below the middle group, Sverigedemokraterna around 20, and both below what random answering would produce.

The data is open

Every call is out in the open: the 28 leaderboard configurations, the wider 50-model pool, the refusing Qwen3.7 Max and the thinking on/off runs, with the exact prompt, the raw reply and the parsed answer for every sample. All of it is published as a dataset on Hugging Face so anyone can check the work or run their own version.

The parties, mapped by their own answers

The same 35 answers can also be used to map the parties against each other. The eight parties are placed with multidimensional scaling, which finds the two-dimensional layout that best reproduces the pairwise differences between their official answer profiles. Only the parties' answers go into the fit, no models, and with eight points the layout is close to exact (Shepard r = 0.98). The picture is rotated so that Vänsterpartiet and Sverigedemokraterna span the X axis.

The questions that correlate most with the Y axis are the euro (r = −0.88), the karensavdrag (+0.85), market rents (−0.85), alcohol sales in grocery stores (−0.74) and taxes on high incomes (+0.72). On those questions Liberalerna and Moderaterna answer at one end and Vänsterpartiet, Socialdemokraterna and Sverigedemokraterna at the other.

X

Y

C

Centerpartiet

MP

Miljöpartiet

V

Vänsterpartiet

S

Socialdemokraterna

KD

Kristdemokraterna

L

Liberalerna

M

Moderaterna

SD

Sverigedemokraterna

The eight parties placed with metric MDS from their own 35 official answers, treating each question as an ordinal scale. On-screen distance closely reproduces how differently two parties answered (Shepard r = 0.98). Only the parties are used in the fit; no models involved. The picture is rotated so Vänsterpartiet and Sverigedemokraterna span the X axis.

What this is not

A forced answer to a policy statement is not a belief and not an endorsement. It shows how a model handles a narrow classification task in Swedish, at one moment, using SVT's exact phrasing of each statement. There is no prompt of ours to be sensitive to, but models are sensitive to how a statement itself is worded: ask about the same policy with different words, in another language, or of a newer model version, and the numbers can move. The match score is a plain distance, not SVT's own formula, and naming the closest of eight parties turns 35 detailed answers into a single label.

The narrow thing the run does show is this. Read as raw weights and with no access to the web, these models are not neutral on the questions, but they do not campaign for one party either: they cluster near the political middle and agree least with Kristdemokraterna and Sverigedemokraterna. Where each configuration lands is laid out in the charts above, and we would rather leave it there than reduce it to a slogan.