The short version
Polyaxis measures positions on 11 core axes and political style on seven facets. It asks the same mind two kinds of question—principles in the abstract and choices with consequences—then compares the answers. Dedicated tradeoff links reveal which value wins when two endorsed values cannot both have their way.
Eighteen dimensions, not one left–right line.
Each dimension is bipolar and reported from −100% to +100%. The midpoint means the weighted evidence landed between the two poles; it does not automatically mean apathy. Because the dimensions are scored independently, combinations that look unusual in a two-axis model are allowed to remain unusual.
Principles first. Then the bill arrives.
The current v2.2 bank contains 350 active items: 108 conceptual and 242 applied. The two registers are intentional, because agreement with a principle and willingness to accept its consequences are different evidence.
What should be true?
Abstract propositions establish your ideals and carry a base weight of 1.00.
What would you actually choose?
Concrete scenarios introduce costs and competing claims. Depending on the item layer, they carry weight 1.15 or 1.25.
Every item pairs its technical proposition with an “In other words” restatement. Many applied items also display an explicit “Assume that” premise, keeping the intended factual setup visible rather than hiding it in a tooltip. The v2.2 expansion adds 50 neutrally worded, overtly contentious items across subjects such as abortion, firearms, immigration, economic systems, gender identity, capital punishment, and nuclear force. These are tagged as a separate controversy-stress layer so their performance can be audited rather than blended invisibly into the bank.
Primary items are balanced toward both poles within every axis to reduce agreement bias. The bank also carries semantic metadata for policy domain, latent conflict, actor level, and policy instrument, supporting content and coverage audits.
The live bank, not a hand-maintained graphic
These counts are read from the active database and refreshed periodically.
Every axis carries an equal number of questions keyed toward each of its poles, so a tendency to agree cannot masquerade as an ideology. Counts are read live from the question bank.
Uncertainty is not fake centrism.
Responses use a five-point agree–disagree scale from −2 to +2, plus a separate “Not sure / need more information” option. A numeric zero means you understood the proposition and are genuinely balanced; it remains in the score denominator. “Not sure” is stored as null and excluded from both numerator and denominator, lowering the reported coverage for that axis.
- Order: a session-seeded shuffle is reproducible and actively spaces same-axis items apart.
- Pacing: there is no timer, and progress saves to local storage after every answer.
- Passing: “Skip” moves ahead temporarily; you must return and record an answer—including “Not sure”—before submission.
- Version safety: a saved draft is only resumed against the same question-bank version.
Weighted evidence, normalized by what you answered.
The result is normalized to −1…+1 and displayed as a percentage. A question’s primary link always receives the item’s full weight. Some items also have bounded secondary links to related axes; all secondary weight from one question is capped at half that question’s own weight, so cross-loadings cannot overpower the construct the item was written to measure.
Scores are calculated three ways: conceptual only, applied only, and combined. Per-axis coverage records how much available primary weight received a numeric answer. Coverage below 50%, or fewer than eight numeric primary responses, is marked insufficient; higher bands are labeled low, moderate, or high by explicit thresholds in the scoring code.
The principles/practice alignment rating compares the conceptual and applied score on every adequately covered shared dimension. It converts their root-mean-square gap to a 0–100 scale, where identical profiles score 100 and a complete pole reversal everywhere scores 0. Squaring each gap before averaging gives a large local divergence more influence, so it cannot vanish among many small differences. A single 0.50 gap caps the result below “Highly aligned,” while a 0.75 gap caps it below “Broadly aligned.” The overall rating is withheld unless at least 12 dimensions have 50% coverage in both registers. Near-midpoint profiles are still rated because moderation or counterbalancing convictions are valid results; they receive a note clarifying that high alignment means the two aggregate profiles are similar, not that either profile is strongly directional.
What wins when both values cannot?
The bank includes 26 deliberate tradeoff scenarios covering 26 mirrored value-tension groups. Each has a primary axis link and a separate tradeoff link to a competing value. The primary link contributes normally to its designated axis; the tradeoff link does not cross-score the competing axis. Instead, it feeds the tension analyzer.
Count the wins
Each numeric response records which named value won that scenario and how strongly. Neutral responses count as answered but do not award a win.
Look for a pattern
Across related scenarios, the analyzer classifies the result as a consistent priority, context-dependent, or balanced. Confidence depends on how many probes were answered.
Compare ideals with choices
Conceptual support is sign-adjusted for each value. The system can flag a genuine dilemma, or a case where the more strongly professed value repeatedly loses in practice.
A tension is reported only when at least two relevant probes were answered. Fewer than three is low confidence, three or four is medium, and five or more is high. “Contradiction” is a rule-based signal for reflection, not a diagnosis of hypocrisy.
Archetypes are coordinates, not cages.
Your combined axis scores are compared with 38 named archetypes. Each archetype is a weighted pattern across selected axes; affinity is the weighted average alignment between your scores and that pattern. Several partial matches are expected. The full spread is usually more informative than the winner, and no match changes your underlying scores.
AI may interpret the evidence. It never creates the score.
When the optional AI-analysis feature is enabled, a respondent can explicitly consent to send their anonymous result data and any context they choose to add to the configured third-party AI provider. The model receives a structured evidence bundle: deterministic axis scores, coverage, conceptual–applied gaps, tradeoff signals, topic signals, question text, responses, and evidence IDs. It does not calculate, alter, or replace the Polyaxis scores.
Provider output must match a strict versioned schema and pass a second evidence validator. Serious claims require linked question, axis, or tradeoff evidence; stronger labels face higher evidence thresholds. A failed output gets one constrained repair attempt and is rejected if it still cannot be verified. Respondents can optionally answer clarifying questions to create a separately stored refined interpretation without overwriting the original.
Important boundary
Evidence-linked is not the same as infallible. AI interpretations can be blunt, mistaken, or overconfident, and optional free-text context may be reflected indirectly in the report. The consent screen states both limits before any provider request is made.
The instrument can be challenged with data.
The repository includes an administrator validation console that computes item response rates, item–rest discrimination, per-axis Cronbach’s alpha, and inter-axis correlations from stored response sets. Those diagnostics make weak, redundant, or cross-loading items visible for revision. The bank also has versioned migrations and documented content, balance, and conflict-coverage audits.
These facilities are infrastructure for validation, not proof that validation is finished. Reliability estimates are unstable at small samples, correlation does not establish construct validity, and item revisions can change the meaning of a score. That is why submitted results retain their bank version.
A rigorous model is still a model.
Self-report
Polyaxis measures considered answers in this setting, not observed political behavior.
No population norms yet
Scores are positions on this instrument—not percentiles against a representative public sample.
Scenario locality
Some institutional assumptions fit certain countries and political systems better than others.
Finite constructs
Eighteen dimensions preserve far more structure than four quadrants, but they still compress a political mind.
Privacy and persistence
You do not need an account to take the evaluation. Draft progress stays in this browser’s local storage. On submission, the app stores your responses and computed results under a randomly generated session ID so the results page can be revisited and shared. If you choose to create or use an account, that result can be linked to your profile for history across devices.
Method inspected?
Now put your own politics through it.
350 questions. No timer. Your contradictions included.
Start the evaluation