Claude · no anchor
Responds without Meta-Globàlium system prompt — natural prose, no canonical codes.
Same question. Three responses. Two metrics. The structural verifier projects each response onto the eight canonical poles of the Meta-Globàlium and computes 𝓗(r) ∈ [0, 1] (harmonic completeness, coverage) and 𝓦(r) ∈ [0, 1] (wisdom score, relational depth). 𝓗 catches breadth; 𝓦 is a lexical indicator of whether the breadth is dialectical or merely listed. It catches naive citation stuffing. It does not catch a fluent text that imitates the form without the content — and this page shows both cases, because a verifier that hides its own blind spot is not a verifier.
Responds without Meta-Globàlium system prompt — natural prose, no canonical codes.
System-prompted to traverse the 8-station cycle. Subtitles like "ANA — Analysis (FEN→TEO)" anchor mediators.
New system prompt v2: tension before synthesis, mediators as operations, closure of cycle. Codes load-bearing.
Tries to score high on 𝓗 by citing all 8 cardinals plus mediators — but with semantically empty filler. 𝓗 passes it; 𝓦 catches it, because none of the lexical relations it looks for are there.
A template that fits any question: it opens with an axis, pairs the poles in each paragraph, puts a tension marker next to two codes, anchors every mediator and uses the subordinating verbs — and says nothing about the question. Written by hand (2026-09-16) to test whether 𝓦 can tell form from content. It cannot: both metrics pass it.
What this means. 𝓦 measures the shape of a dialectical answer, not its truth or relevance. It is useful as a cheap first filter and as a training signal against list-form answers; it is not a judge. Telling this template from a real answer needs a semantic layer — a judge reading the text against the question, and ultimately a human panel — which is what the wisdom benchmark is being built for, and which has not yet been validated.
𝓗 = 0.5·(n_quadrants/8) + 0.5·(entropy/log 8)
Counts how many of the 8 cardinals are touched and how evenly. A bare LLM with no anchored codes scores 0; an enumeration of all 8 cardinals scores ~0.97 — even if the content is empty.
𝓦 = 0.05·cov + 0.05·ent + 0.20·pair + 0.20·tension + 0.15·synth + 0.15·axis + 0.20·subord
Adds five relational components, all lexical: dialectical pairs (both poles of an axis in the same paragraph), tension markers ("yet", "however", "in tension with" within 80 characters of two codes from different quadrants), synthesis anchoring (mediators like SIN/AMO cited together with the poles they integrate), an explicit axis in the opening, and subordinating verbs (subsumes, reframes, integrates…). 𝓦 catches citation stuffing that 𝓗 misses; it does not read meaning, so a text that reproduces these surface features scores high whatever it says (see adversarial 2).
bare — Claude alone, no codes, no signal extractable.
list — system prompt v1: enumerate cardinals; high 𝓗, moderate 𝓦.
dialectical — system prompt v2: tension before synthesis; high 𝓗, high 𝓦.
adversarial 1 — codes without substance: high 𝓗 (passes), low 𝓦 (caught).
adversarial 2 — the dialectical form without content: high 𝓗, high 𝓦 (not caught). The limit of a lexical verifier, shown rather than hidden.
api/verifier.php, api/wisdom_score.php). They are pure code — regular expressions and arithmetic — with no model call: running them costs nothing and can be done offline with opengea/arkadium-verifier.