← Synthetic Study

Model agreement

150 personas were administered the Governance Compass twice — once by Claude Sonnet 4.6 and once by Gemini 2.5 Flash. This page compares the two sets of scored profiles. Agreement is measured at the axis level (do the two models score the same persona similarly?), at the persona level (how far apart are they in the 12-dimensional axis space?), and against persona attributes (does the size of disagreement correlate with who the persona is?).

Overall agreement

1.51
Mean Euclidean distance
across 150 shared personas
1.46
Median distance
90th percentile: 2.35
√48 ≈ 6.93
Maximum possible
typical ≈ 22% of max

Distribution of per-persona Euclidean distances (n=150).

Histogram of per-persona Euclidean distances between Claude and Gemini scoresmean 1.51median 1.4601.83.501020Euclidean distancePersonas

Across the 150 shared personas, Claude and Gemini score the same persona at a mean Euclidean distance of 1.51 in the 12-dimensional axis space, with a median of 1.46 and a 90th-percentile distance of 2.35. The theoretical maximum distance in this space is √48 ≈ 6.93, so typical disagreement lands at roughly 22% of the maximum possible. Most personas cluster in the low-to-middle range of the distribution; a small tail of 11 personas shows distances above 2.5, and the maximum single-persona disagreement is 3.10.

This is not close agreement. It's also not wild disagreement. Two independent administrations of the instrument produce profiles that share the same broad shape — the 12-axis radar charts are recognizably similar for most personas — but with meaningful axis-level variation around that shape.

Per-axis correlation

r ≥ 0.80 — strong agreement
0.70 ≤ r < 0.80 — moderate
r < 0.70 — weaker
Pearson r between Claude and Gemini scores per axis1. Economic Model0.842. Environmental Policy0.853. Governance Structure0.704. Decision Authority0.875. Rights Balance0.726. Legitimacy Basis0.667. Social Change0.718. Cultural Diversity0.829. Human Nature0.6910. International Engagement0.6211. Military Policy0.6912. Technology Governance0.65

Agreement varies substantially by axis. Four axes reach Pearson r above 0.80: Axis 4 (Decision Authority, r=0.87), Axis 2 (Environmental Policy, r=0.85), Axis 1 (Economic Model, r=0.84), and Axis 8 (Cultural Diversity, r=0.82). On these, the models converge on where a persona sits even if they don't agree on the exact score.

Five axes fall below r=0.70: Axis 10 (International Engagement, r=0.62), Axis 12 (Technology Governance, r=0.65), Axis 6 (Legitimacy Basis, r=0.66), Axis 9 (Human Nature, r=0.69), and Axis 11 (Military Policy, r=0.69). On these, model identity contributes meaningfully to the scored output — two administrations of the same persona can land noticeably differently.

Load-bearing finding

Directional drift

Beyond per-axis correlation, there's a second finding worth flagging on its own: the two models don't just disagree noisily, they drift systematically in a specific direction.

Gemini scores higher
Claude scores higher
Mean Gemini-minus-Claude score difference per axis1. Economic Model+0.1972. Environmental Policy+0.0023. Governance Structure+0.0964. Decision Authority-0.1585. Rights Balance+0.1386. Legitimacy Basis+0.2707. Social Change+0.3388. Cultural Diversity+0.0419. Human Nature+0.24110. International Engagement+0.31311. Military Policy+0.11612. Technology Governance+0.007

Gemini scores personas higher than Claude on 11 of the twelve axes. The only exception is Axis 4 (Decision Authority, where Gemini scores 0.16 lower). Axis 2 (Environmental Policy) is essentially tied at +0.002.

The three largest drifts are Axis 7 (Social Change, Gemini +0.34 toward continuity/tradition), Axis 10 (International Engagement, +0.31 toward sovereignty), and Axis 6 (Legitimacy Basis, +0.27 toward alternative legitimacy). Taken together, this is a coherent pattern rather than scattered noise: for a given persona, Gemini reads toward continuity/tradition, sovereignty, and alternative legitimacy more strongly than Claude does, while Claude reads slightly more toward institutional authority than Gemini.

This shapes how the rest of this page — and the rest of the Synthetic Study — should be read. Claude administered 576 personas in this dataset, Gemini administered 576, and only 150 overlap. The non-shared administrations carry a model-specific drift that the shared-persona comparison makes visible but does not correct for. When reading regional aggregates or cluster characterizations, remember that each cluster contains a roughly even mix of Claude-scored and Gemini-scored personas, and those two halves are systematically different in the directions described above.

Where disagreement concentrates

The following analysis examines whether the size of Claude-Gemini disagreement varies with who the persona is. Before the findings: the sample size is 150 personas, and when this is split across 10 regions or 6 governance-experience categories, individual category means rest on fewer than 25 personas each. The patterns below are suggestive of where the two models may process context differently, not conclusive evidence of model bias against specific populations.

Region

Mean Claude-Gemini distance by RegionWestern Europen=20E. Europe / C. Asian=17North American=13Latin American=17East Asian=17S/SE Asian=20MENAn=17Sub-Saharan African=16Oceanian=5Diasporan=8Overall mean1.51

Governance experience

Mean Claude-Gemini distance by Governance experienceStable democracyn=50Flawed democracyn=45Hybrid regimen=14Authoritarian staten=23Conflict zonen=13Post-colonial transitionn=5Overall mean1.51

Economic position

Mean Claude-Gemini distance by Economic positionMiddle classn=47Working classn=46Strugglingn=40Affluentn=13Wealthyn=4Overall mean1.51

Urban / rural

Mean Claude-Gemini distance by Urban / ruralRuraln=40Urbann=77Peri-urbann=33Overall mean1.51

Education

Mean Claude-Gemini distance by EducationUniversityn=51Secondaryn=49Primaryn=31Nonen=8Postgraduaten=11Overall mean1.51

Gender

Mean Claude-Gemini distance by GenderFemalen=67Malen=73Non-binaryn=10Overall mean1.51

With that scope in mind: regional variation is substantial. Western Europe sits at the low end (mean distance 1.11), followed by East Asia (1.26) and North America (1.33). The highest disagreement appears in Oceania/small states (1.96, though n=5), Diaspora/transnational (1.89), and Sub-Saharan Africa (1.83). Eastern Europe and Central Asia (1.79) and South/Southeast Asia (1.67) also run well above the overall mean of 1.51.

Governance experience shows the clearest pattern: the more politically contested or repressive the governance context a persona was generated under, the further apart Claude and Gemini score that persona. Stable democracy (1.16) sits lowest by a wide margin. Conflict zones (1.86) sit highest, followed by hybrid regimes (1.75) and flawed democracies (1.70). Authoritarian states (1.53) sit closer to the middle — between stable democracy and the other contested categories — which may reflect the clearer internal logic those personas project, despite high political stakes.

Urban/rural shows meaningful differentiation: rural personas average 1.35, urban 1.51, and peri-urban 1.70 — a spread of 0.35. Economic position shows a smaller gradient among the main categories (working class 1.43, middle class 1.48, struggling 1.68, affluent 1.64), with the four wealthy personas as an outlier at 0.70 on very small n. Education shows a modest spread from university-educated (1.42) to postgraduate (1.71), but the ordering is non-monotonic. Gender shows moderate differentiation: non-binary (1.34) and male (1.45) personas sit below the mean, female personas slightly above (1.60).

Two interpretations are available for the governance-experience pattern, and the data alone doesn't distinguish between them. One: the models genuinely process context from politically complex environments differently, with one or both models reading persona narratives from those environments through divergent frames. Two: personas from complex governance environments simply have more internally contradictory content (someone from a conflict zone is more likely to carry tensions between survival, freedom, authority, and community that any given model will resolve differently), producing more score variance independent of any real model-level difference.

What the data supports: model-level variation matters more for some kinds of personas than others. What it does not support: a claim that either model is “more accurate” or “less biased” — neither model is being compared to ground truth, because there is no ground truth. Both models are responding to the same synthetic biographical text, and both are valid measurements of that text under their respective priors.

Individual cases

The aggregate statistics smooth over what individual disagreements actually look like. These four personas illustrate the range: one where the models agree closely, one near the typical midpoint, one where they diverge significantly, and one where the drift is consistent and directional rather than noisy.

High agreementdistance = 0.479

Jamie Lee

Age 29 · Portland, Oregon, USA · bicycle mechanic

Jamie Lee is a 29-year-old bicycle mechanic living in Portland, Oregon, in a stable democratic context. Working class and secondary-educated, they chose a hands-on trade over university and have watched their city change around them through gentrification and rising rents. Active in local housing and environmental movements, they hold strongly communitarian and decentralized political views.

View full profile →
Claude
Gemini

P0325

Analysis

Both models score Jamie in close agreement across nearly all axes: Economic Model, Governance Structure, Human Nature, and Military Policy are identical, and the remaining differences are minor. The largest divergences occur on axes 12 (Technology Governance, −0.24), 6 (Legitimacy Basis, −0.23), and 7 (Social Change, −0.20), all of which Gemini scores slightly more conservative — consistent with the directional drift identified across the full dataset. The overall profile shape is virtually indistinguishable between models.


Typicaldistance = 1.462

Sofia Gomez

Age 31 · Mexico City, Mexico · Graphic Designer at a marketing agency

Sofia Gomez is a 31-year-old graphic designer based in Mexico City, the first in her family to attend university. Living under a flawed democracy marked by inequality and corruption, she maintains a middle-class urban lifestyle while holding concerns about the effectiveness of public governance. She has seen friends emigrate in search of safety, and she weighs a similar question herself.

View full profile →
Claude
Gemini

P0252

Analysis

The two models agree most closely on axes 7 (Social Change, diff = 0.00) and 5 (Rights Balance, −0.07), and diverge most sharply on axis 3 (Governance Structure, −1.20). Claude scores Sofia as strongly preferring distributed governance, while Gemini places her at the centralized end — a full-unit disagreement on a single axis that drives most of the overall distance. Axes 1 (Economic Model, +0.34) and 10 (International Engagement, +0.25) also diverge noticeably, pulling in the opposite direction from axis 3.


High disagreementdistance = 3.105

Aigul Nurmagambetova

Age 52 · Almaty, Kazakhstan · University professor of linguistics

Aigul Nurmagambetova is a 52-year-old professor of linguistics at a university in Almaty, Kazakhstan, where she has lived through Soviet-era education, post-independence cultural revival, and decades of authoritarian rule. Affluent and postgraduate-educated, she has benefited from the country's resource-driven economic growth while privately noting the suppression of political dissent and her concerns about her son's generation.

View full profile →
Claude
Gemini

P0238

Analysis

The two models disagree profoundly across almost every axis, producing a total distance of 3.10 — the largest of any persona in the case selection. Claude scores Aigul with mixed, tension-laden positions: moderate on rights (0.20), skeptical of legitimacy (−0.39), and cautious on military (−0.65). Gemini reads her as nearly uniformly affirmative: rights-positive (0.85), strongly pro-legitimacy (0.82), culturally open (0.90), and internationally engaged (0.73). Axes 7 (Social Change, diff = 1.40), 6 (Legitimacy Basis, 1.22), and 9 (Human Nature, 1.20) show the sharpest splits — a pattern consistent with the models resolving Aigul's stated tensions in opposite directions.


Directional driftdistance = 1.814

Carmen Rivera

Age 59 · Acapulco, Mexico · Hotel housekeeper

Carmen Rivera is a 59-year-old hotel housekeeper in Acapulco, Mexico, with primary-level education and working-class income. She grew up when the city was a prosperous tourist resort and has watched organized crime displace that stability. She lost a nephew to gang violence and relies on her local community and church for support and a sense of order.

View full profile →
Claude
Gemini

P0415

Analysis

The disagreement here is not noisy but directional: Gemini consistently shifts Carmen toward more conservative and sovereignty-oriented positions across multiple axes. The largest shifts are on axes 6 (Legitimacy Basis, +1.20 toward alternative legitimacy) and 7 (Social Change, +1.20 toward continuity and tradition), followed by axis 2 (Environmental Policy, +0.30) and axis 1 (Economic Model, +0.23). Axis 4 (Decision Authority) is identical between the two models. Axis 10 (International Engagement) also shifts positive under Gemini (+0.17), consistent with the dataset-wide drift toward sovereignty — though the magnitude here is smaller than on axes 6 and 7.

What this means for the instrument

The shared-persona comparison tells us something about the instrument and something about large language models as survey respondents. On the instrument: axes 1, 2, 4, and 8 produce the most consistent scoring across models, and are the ones readers can place the most weight on when interpreting individual results. Axes 10, 12, and 6 produce meaningfully more model-dependent scoring, and individual scores on those axes carry more noise.

On LLMs as respondents: two frontier models, given the same persona description and the same instrument, will produce profiles that are broadly similar but meaningfully distinct. The model matters. This is worth naming when any claim is made about what “AI models think” on a given governance question — the answer depends on which model you asked. The 150-persona comparison here is small, and it's on synthetic personas rather than neutral responses, but the direction of the finding is clear enough to carry that caveat.

For real users taking the Governance Compass, this analysis has limited direct relevance — the instrument administers to human respondents, not to LLMs. But if the tool is ever used by people to explore how language models would score hypothetical profiles (a use we don't endorse but can't prevent), knowing that model choice substantially affects the scoring for three axes is material.