Background Public prediction records are hard to evaluate: posts can be edited or deleted, and even honest records rarely separate anteriority (did this exact claim exist unchanged before the event?) from skill (do the claims carry predictive information?). A motivating case: a public forecast naming mass-casualty attacks on public gatherings in Moscow in a March–April 2024 window was published and server-timestamped on 4 September 2023 — ~200 days before the Crocus City Hall attack and 185 days before a comparable official alert. Methods Every forecast was published verbatim before its window; each text was canonicalised, SHA-256-hashed into a manifest, and anchored to the Bitcoin blockchain via OpenTimestamps — the anchor fixes integrity (the bytes unchanged since the anchoring block), while anteriority rests on platform-issued timestamps, pre-event public distribution, and third-party archives (§2.1.2); a four-level rubric (HIT/NEAR/PARTIAL/MISS) was fixed at seal time and misses retained. We formalise a five-step source-agnostic scoring protocol for sources whose generative mechanism is unavailable or unexplained: demand cryptographic anteriority; freeze the claim verbatim; grade the stream, never the anecdote; set weight by calibration, not theory; and let the frozen denominator filter uninformative sources. Results Across 92 graded forecasts (73 HIT, 6 NEAR, 4 PARTIAL, 9 MISS) the self-assigned aggregate Brier is 0.0958 (launch subset 0.036; strategic warning 0.116). On this small, high-confidence sample a base-rate baseline ties the aggregate Brier — conceded up front. A pre-specified clustered luck test — correlated calls collapsed into independent events, strict all-HIT scoring, luck prior floored at 0.5 per event — yields 51 of 68 strict successes (exact binomial p = 2.2 × 10−5; break-even luck prior ≈ 0.65). A five-way calibration-integrity suite is reported with sample-size ceilings. All statistics recompute from public artifacts with zero-dependency tooling. Conclusions The contribution is the protocol — belief-independent evaluation of unexplainable forecast sources — not validation of the disclosed generative method (Vedic jyotish, treated as a black box). Applications to warning analysis and war-game parameterization are outlined; the limitations — self-grading, sample size, operator selection — are load-bearing.
On 7 March 2024, the United States Embassy in Moscow issued a public security alert stating that it was monitoring reports that extremists had imminent plans to target large gatherings in Moscow, including concerts, and advising citizens to avoid large gatherings for the next 48 hours.1 Fifteen days later, on the evening of 22 March 2024, gunmen attacked the Crocus City Hall concert venue in Krasnogorsk, on the northwestern edge of Moscow, killing 145 people.2 This paper makes no claim regarding the attack’s perpetrators or sponsors, and grades nothing about attribution.
We want to state plainly how that alert should be read: as a benchmark of excellence. A public, pre-event, correctly-targeted alert — right city, right target class, right threat type — issued ahead of a mass-casualty attack is among the rarest artifacts in the open record of warning analysis, and fifteen days of public lead is, for structural reasons developed in §1.2, close to the ceiling of what collection-driven warning can deliver against an attack of this kind. Nothing in this paper is a criticism of that alert. Our subject is a second, earlier timeline to the same evening.
On 4 September 2023 — 200 days before the attack, and 185 days before the embassy alert — the thirty-seventh post of a numbered 59-post public thread read, verbatim:
“Terrorist attacks on public gatherings are possible in Moscow in March and April in an attempt to sow panic amongst the general public of Russia, who have been well-insulated from the harsh realities of this conflict.”3
City; target class; two-month window; stated intent — fixed in advance, in one sentence, in public. The author of that sentence is the author of this paper, and the author of a forecast is the least persuasive witness to its date: a forecaster exhibiting his own hit, after the event, from his own archive, is a pattern that analysts and referees are correctly trained to discount. The paper therefore asks the reader to take nothing on testimony, including that sentence. Its text is committed by SHA-256 hash into a public manifest anchored in the Bitcoin blockchain; its publication timestamp was issued by a platform server the author does not control and decodes from the post identifier by pure arithmetic (§2.1.2); and the full verification procedure — executable offline, on an air-gapped machine, in roughly a minute of machine time — is specified in §2.2 before any scoring is reported.
One further condition governs everything in this paper. Every forecast discussed here was unsolicited. The forecasts were posted publicly; in several documented, dated instances they were publicly addressed to officials (§3.7). No recipient ever acknowledged or responded to any of them, and the record claims nothing beyond the act of publication itself: nothing in this record was, to the author’s knowledge, ever read, used, or acted upon by anyone, and this paper makes no claim otherwise. The record exists in public; that is the entire claim.
The Crocus timeline is the motivating case, not the contribution. The contribution is a protocol: a procedure by which an evaluator — a referee, an analyst, a scientist — can extract measurable signal from a forecast source whose generative mechanism is unavailable, implausible, or unexplained, without belief in the source ever becoming a dependency of the analysis (§2.5). The sealed record examined here, 92 graded forecasts with the misses retained, is the protocol’s worked instance.
The gap between the two timelines is not a difference of degree, and it is worth stating structurally, because it defines what kind of evidence a 200-day lead can and cannot be.
Collection-driven warning works backward from a plot in being. Someone must decide, recruit, communicate, move money, travel, rehearse — and each of those acts sheds signals that collection can, with skill, intercept. This mechanism has a structural property that no resources can remove: the warning clock cannot start before the plot clock starts. There is nothing to collect against a plan that does not yet exist. Attacks on soft targets are, in the documented pattern, short-cycle plans: long timelines multiply exposure and interception risk for operations whose logic is cheapness and speed. It follows that at 200 days before a short-cycle attack, the expected quantity of collectible plot signal approaches zero. We are careful here: we claim no knowledge of when the Crocus plot began, and the argument does not depend on the case. The point is structural — a fifteen-day public lead is excellent work near a physical ceiling of the collection paradigm.
A forecast sealed 200 days out, then — whatever one thinks of its source — is evidence about a different variable. On the argument above it could not have been reading a plot; there was plausibly no plot to read. What its own text named was a condition of the strategic landscape and a pressure acting on it: a public “well-insulated from the harsh realities of this conflict,” and an “attempt to sow panic,” against public gatherings, in a specific city, in a specific two-month window. Whether any method can systematically read such a variable is precisely the kind of question that cannot be answered by one case, which is why this paper’s unit of analysis is never the single call. It is the scored stream (§5.1).
A forecast is only evidence if it can be shown to have existed, in its claimed form, before the event it forecasts. Most public prediction records fail this test. Social-media posts can be edited or deleted; “I called it” claims are made after the fact with no contemporaneous artifact; ambiguous wording lets a forecaster claim credit for almost any outcome (the hindsight, vagueness, and “I-was-basically-right” biases quantified in elite forecasters by Tetlock and colleagues4). Even an honest record is usually un-auditable: no third-party-verifiable timestamp, no fixed scoring rule, no guarantee failures were retained beside successes.
These are two distinct problems, usually conflated:
1. Anteriority and integrity — did this exact claim exist before the event, unchanged since? A cryptographic and archival question, independent of whether the claim was any good.
2. Skill — given that the claims are genuinely anterior, do they carry predictive information? A statistical question, answered with proper scoring rules5,6 and calibration analysis.
A credible record must dispose of (1) before (2) is meaningful, because without (1) any apparent skill could be an artifact of post-hoc editing or selective recall.
There is a third problem, posed to any evaluator confronted with a source that appears to perform but whose mechanism is unavailable or unexplained: what procedure extracts signal from such a source without the evaluator’s belief becoming a dependency? This is not an exotic question. It is the standing epistemic position of any analyst handling an unsolicited warning — an anonymous letter, an unattributed tip, a walk-in source — and, more generally, of any scientist confronted with an anomalous claim attached to a well-kept record. The forecasting-tournament literature4,7 answers how to score forecasters who volunteer for tournaments; it does not specify a procedure for the source that arrives uninvited, with an implausible mechanism and a dated archive. Section 2.5 specifies one.
This paper describes a protocol making the three questions independently checkable, and applies it to a live corpus:
• An anteriority-and-integrity layer. Every forecast’s exact text is published to public infrastructure before its event window, canonicalised, SHA-256-hashed into a manifest, and the manifest digest committed to the Bitcoin blockchain via OpenTimestamps,8 in the lineage of cryptographic timestamping introduced by Haber and Stornetta.9 Anteriority — that the sealed text existed before the event — rests on the platform publication timestamp and the post’s public distribution before the event, corroborated by independent third-party archives; the Bitcoin anchor establishes integrity: the manifest bytes have been unchanged since the block that commits them. For forecasts sealed and anchored before their event windows closed, the anchor’s block time additionally precedes the outcome and the two properties collapse into a single proof (§2.1.2).
• A falsifiability layer. A grading rubric is pre-committed at seal time. Outcomes are graded against the sealed claim, four-level, with misses retained in the denominator, grades frozen in a separately anchored ledger.
• A scoring layer. Proper scoring rules (Brier, log-loss), IPCC-AR6-style calibration bins,10 Information Yield — an information-theoretic statistic measuring against-consensus surprise per call — and a five-way calibration-integrity suite (§2.3.5) importing the standard model-validation instruments of clinical prediction and actuarial science.11,12
• A reproducibility layer. A zero-dependency public verifier and a CC-BY corpus let anyone recompute every hash, the time anchor, the Brier score, and every reported statistic from frozen inputs, removing the author from the trust chain (§2.2; Data availability).
• A distribution-and-delivery layer. Because the forecasts were published on a platform with a queryable archive, a third party can reconstruct not only that a claim was anterior but how widely it was distributed before the event and to whom it was publicly delivered — both from platform-native metadata, independent of the author (§3.7).
• A source-agnostic scoring protocol. The methodological headline of this version: a five-step procedure (§2.5) by which an evaluator can score an unexplainable forecast source — demand anteriority, freeze the verbatim claim, grade the stream, weight by calibration, and let the frozen denominator filter uninformative sources — with belief in the source never entering the procedure at any step.
• Applications. An applications section (§4) develops what a belief-independent scored stream offers two analytic settings — warning analysis and war-game parameterization — stated as offers of tradecraft with no counterfactual claims attached.
We are explicit that this is a protocol contribution. The scoring is self-assigned and the corpus is small; nothing here is an independently validated skill claim, nor a claim about the validity of the underlying generative method. The threats-to-validity section (§6) develops these limitations in detail and is, deliberately, the most load-bearing section of the paper.
2.1.1 Canonical string and hash
Each forecast object reduces to a canonical four-field string objectId|dateIssued|title|claim, where objectId is the identifier of the public post carrying the verbatim forecast (an X status ID or YouTube video ID), dateIssued the publication date, title a human-readable label, and claim the full sealed text. This string is SHA-256-hashed to a per-record digest; all digests collect into a manifest (seal-manifest.json) with its own manifestHash. The pipe-delimited template is published in the manifest header (hashInputTemplate), so canonicalisation is not secret: any party can reconstruct the preimage from published fields and confirm the hash.
2.1.2 Anteriority versus integrity — two properties, two clocks
The seal provides two properties on different evidence. Integrity (the text is unchanged since sealing) rests on the SHA-256 digest and the OpenTimestamps Bitcoin anchor: any edit changes the hash, and the manifest bytes have been unchanged since the block that commits them. Anteriority (the text existed before the event) rests on the platform publication timestamp and the post’s public distribution before the event, corroborated by independent third-party archives; a property specific to X makes the platform clock checkable offline — the post’s numeric status ID encodes its server-side creation time. The Twitter/X “snowflake” identifier satisfies.
timestamp_ms = (id >> 22) + 1288834974657
so the creation instant decodes from the ID by pure arithmetic, with no API call and no trust in the author. Every advisory in this corpus publishes its status ID, so a reviewer can confirm each seal’s date offline. For the Crocus advisory, status ID 1698760248033476914 decodes to 4 September 2023, 18:09:39 UTC.
An honest scope statement belongs here, because a proof overstated is a proof destroyed. A cryptographic seal proves integrity and anteriority. It proves nothing — nothing — about correctness: a sealed forecast can be sealed, anchored, timestamped, and wrong. The seal converts “trust me” into “check me”; it does not convert a claim into a fact. The two properties also ride on two different clocks with two different failure modes, and we name the seam rather than leave it to a referee. For any forecast sealed and Bitcoin-anchored before its window closed, the anchor precedes the outcome, so integrity and anteriority collapse into one proof. For events that resolved before the full-corpus manifest was anchored (Crocus among them), anteriority rests instead on the platform clock decoded above, on the post’s documented public distribution before the event (the Crocus post had accumulated 653,566 views on the platform’s public counter before 22 March 2024; §3.7), and on independent third-party archives. Two clocks, two independent failure modes; fabricating the record would require defeating both, in public, against preserved copies.
EXHIBIT 1· THE LAYERED SEAL ARCHITECTURE — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 1. Independent verification layers over each sealed call: platform timestamp (anteriority), SHA-256 manifest, OpenTimestamps Bitcoin anchor (integrity), frozen grading ledger, zero-dependency verifier. Rendered live in the online edition (see pointer above).
2.1.3 The manifest and the two anchors
The corpus maintains two independently Bitcoin-anchored artifacts: the seal manifest (which freezes the verbatim claims; currently 103 sealed records) and a separate grading ledger (grading-ledger.json), which freezes the grade and the sealed probability of each graded call (currently 92 entries) and computes the Brier score as plain arithmetic over the frozen terms.13 The two are anchored separately and re-stamped independently, so a change to grading cannot silently alter the seal record and vice versa. The manifest holds more records than the graded ledger because sealed-but-deliberately-ungraded items (prevention-class, ethically excluded, and unfalsifiable-by-design calls; §6.8) are retained in the manifest but held out of the scored denominator, with their exclusion reasons published.
Alongside the 103 manifest records the corpus publishes a separate array of 83 companion seals — corroborating sibling posts drawn from the same sealed threads, each carrying its own SHA-256 digest over the identical canonical recipe and riding the same OpenTimestamps file anchor as its parent record, but deliberately held out of the graded count. Their purpose is provenance, not score: they let a reader confirm that a graded advisory did not stand alone but sat inside a larger, contemporaneously sealed body of dated claims, while their exclusion from the denominator ensures the graded headline (92) can never be read as padded by sibling posts.
A structural consequence deserves emphasis: an anchored record accumulates misses instead of shedding them, because the anchor makes deletion detectable. A fabricator’s first instinct is to prune; the OpenTimestamps proof turns every pruning into a broken hash. Once a manifest is committed into a blockchain, misses cease to be a liability the record’s owner manages and become the exhibit that proves the hits were not curated. In a precise sense the misses are the most evidentially valuable entries on the ledger.
The paper is designed to be checked, not believed, and the check precedes the results. Two zero-dependency commands (Node ≥18, no installation, offline-capable) reproduce the whole paper:
node verify-jyotint.mjs --manifest seal-manifest.json --ots seal-manifest.json.ots # INTEGRITY: recomputes every SHA-256 seal, the Merkle tree (103 inclusion proofs), # the OpenTimestamps Bitcoin anchor, and the Brier from the frozen grading ledger; # exits non-zero on any drift. node reproduce-paper.mjs # STATISTICS: re-derives every reported statistic — the five-way calibration suite, # the clustered luck test, Information Yield, SITA, and the novel axes — from the raw # graded record and checks each against the published value, printing PASS/FAIL per line.
The first command proves the record is intact — every hash, every Merkle inclusion proof, and the Bitcoin anchor recomputed — while each seal date is established on the second clock by the platform-identifier arithmetic of §2.1.2; the second command proves the analysis is an honest projection of the record — no number in this paper was entered by hand. A single-record check is available for any advisory (e.g. --id IA-RU-008 verifies the Crocus entry alone). For fully offline review, a self-contained air-gapped capsule (the manifest, its OpenTimestamps proof, and the verifier, with no external dependencies; Data availability) can be copied to a machine with no network connection; nothing in the verification requires one. The verifier script is short enough to be read line by line before execution, which we encourage. A reviewer who completes these steps has established every date that matters independently of the author: from that point the paper’s phrase “sealed 4 September 2023” is not testimony but a reference to arithmetic the reviewer has already performed.
2.3.1 The four-level rubric
Each closed call is graded HIT/NEAR/PARTIAL/MISS against the sealed text, using a four-dimension conjunction (actor · timing · target · effect/mechanism). A call that names four axes and matches three is not a HIT; NEAR and PARTIAL capture graded shortfalls, and MISS is retained at full weight. The rubric is fixed at seal time and frozen in the ledger. The published grading instruction operates against the grader’s interest: score the sealed pre-event claim text, never the descriptive catalog title assembled afterward. Outcome values are HIT = 1, NEAR = PARTIAL = 0.5, MISS = 0; each entry’s Brier term is the squared gap between the frozen probability and the outcome value, applied identically to hits and misses.13 ( Figure 2.)
Authored vector plate (print edition), from the vector figure set of the companion monograph (CC-BY 4.0; Data availability).
2.3.2 Sealed probabilities
Sealed probabilities are mapped from the linguistic confidence of each claim onto the IPCC AR6 calibrated-likelihood bands10 (a flat declarative → “very likely” 0.90; “likely/expected” → 0.78; “possible/could” → 0.60; and so on). The mapping is applied identically to hits and misses and is published with the calibration artifact.
2.3.3 Information Yield and SITA
Information Yield measures bits of surprise-if-true per call, log2(1/(1 − p_consensus)), from operator-assigned, published, conservatively floored base rates; a call that merely follows the base rate scores zero bits by construction. Priors are contestable and re-derivable by any reader; the statistic is checkable, not oracular. SITA profiles every call on four axes — Specificity and Improbability computed from published inputs; impacT and Actionability judged against a published rubric with transparent weights (S 0.20 · I 0.20 · T 0.30 · A 0.30). The per-axis profile is the primary artifact; the composite is only a summary. Neither statistic enters the Brier score.
2.3.4 The corpus-level luck test
Because many calls within a theatre are correlated (a mega-thread’s sub-claims share a world-state), a naive per-call binomial overstates independence. The luck test therefore clusters correlated calls into independent events, scores each event strictly (any NEAR/PARTIAL/MISS fails the whole event), and floors the per-event luck prior at a coin flip (0.5) — deliberately hostile choices, each operating against the record. The clusters, the sensitivity band, and the floor are published in a live artifact (/api/v1/luck-test.json) and recompute from the graded corpus; a reader who disagrees with any clustering choice can re-partition and re-run.
2.3.5 The five-way calibration-integrity suite
A single Brier number is a summary; it hides why a record scores as it does. Clinical prediction and actuarial science do not accept a scalar in isolation — they demand a validation suite that separates calibration (are the probabilities honest?) from discrimination/resolution (do they separate outcomes?) and reports each with an uncertainty ceiling. We import that discipline wholesale: the Murphy decomposition,11 Spiegelhalter’s Z,12 Cox calibration regression,14 sharpness, and calibration-in-the-large, each computed on the frozen ledger and each travelling with its own stated limitation (§3.5).
2.3.6 The squared-error guardrail
The choice of the Brier score is not cosmetic; it is the property that makes this record hard to game. The Brier score is strictly proper5,6: it is minimised in expectation only by reporting one’s true probability — and, in the limit, by being right. The asymmetry is stark at the confidence this record uses most: a call sealed at 0.78 that resolves HIT contributes (0.78–1)2 = 0.0484; the same call resolving MISS contributes (0.78–0)2 = 0.6084 — the miss costs ≈ 12.6× the hit. This inverts the volume strategy common to public prediction records (emit many calls, foreground the landers, let failures decay off the feed): under a squared, miss-retaining rule, every wrong confident call detonates in the denominator, and no quantity of easy hits dilutes it faster than misses accumulate. The rule rewards exactly one behaviour — sealing fewer, better-calibrated, actually-correct calls. We do not claim the rule proves skill (on this sample a base-rate baseline ties the aggregate Brier; §3.1); the guardrail governs how the record can be gamed, not whether it demonstrates calibration.
2.4.1 The sitting as the unit of production
The natural unit of output in this record is not the individual post but the sitting — a single dated session emitting a numbered mega-thread whose sub-claims are all sealed within the same minutes. Four such sittings in the Russia–Ukraine theatre produced threads of 23, 31, 25, and 59 numbered sub-forecasts; two India sittings (§3.9, §3.10) complete the set of six multi-call seal dates. The sitting matters methodologically for two reasons. First, correlation: sub-claims of one sitting share a world-state, which is why the luck test clusters them (§2.3.4). Second, anti-curation: a sitting is frozen whole — every numbered post server-timestamped within the same minutes, strong claims and weak under the same hash discipline — so its failures cannot be quietly divorced from its successes. The 4 September 2023 sitting, worked in full at §3.8, is the paper’s clearest exhibit of both properties.
2.4.2 The full-corpus recovery
The record’s largest theatre, the Russia–Ukraine war, was forecast in public from February 2022. In July 2026 the complete posting record behind the graded advisories was reconstructed from the platform archive to test a question the graded ledger alone cannot answer: were these forecasts distributed before their events, and to whom? The recovery used the platform’s full-archive search across the operator’s accounts, snowflake decoding (§2.1.2) to confirm each post’s creation instant offline, and the platform’s public_metrics field to recover impression and engagement counts. The recovered set comprises 4,547 Russia-related posts across four accounts. All distribution figures are platform-native and re-pullable by any party from the post IDs in the public corpus; the platform records impression counts only from December 2022 onward, so reach totals are floors.
2.4.3 The corpus funnel
To answer the selection objection at the corpus level, every original owned Russia-topic post was classified into exactly one disposition bucket (374 retweets — amplification of others’ posts, not original forecasts — are excluded upstream), and the classification is published as a static artifact. The funnel is reported as a result at §3.9.3.
2.4.4 Exclusion policy
Whole prediction classes are deliberately held out of the graded denominator with published reasons: named-individual personal-safety forecasts (prevention-paradox unfalsifiability and ethics), live humanitarian catastrophes (ethics, including a mechanically qualifying would-be HIT declined; §6.8), and advisory-register posts that carry no dated falsifiable claim. These exclusions remove apparently favourable material as well as unfavourable; §6.8 argues this is the point.
EXHIBIT 2· CALLS THE SCORE REFUSES — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 2. The exclusion register operating against interest: displayed misses that never counted for leniency, and flattering wins declined because they carried no information. Rendered live in the online edition (see pointer above).
This section states the paper’s headline methodological contribution: a procedure by which an evaluator can extract measurable signal from a forecast source whose generative mechanism is unavailable, implausible, or unexplained, without belief in the source entering the procedure at any step.
2.5.1 The problem of the unexplainable source
The generative method behind the corpus under study is Vedic jyotish — the Indian astrological tradition, practised as a forecasting discipline. It is named once, plainly, and this paper offers no mechanism, makes no claim about why it should work, and does not advocate for the tradition. The evaluative standard adopted is the empirical one: a method earns exactly the credence its sealed, graded, public record earns — no more because a tradition is old, no less because it is stigmatised. Everything in this paper stands or falls on that standard alone.
Having named the method, we deliberately seal it into a black box for the remainder of the paper — not as a rhetorical evasion but because the black box is the methodological point. An unexplainable source is the standing epistemic position of every analyst who has handled an unsolicited warning: a walk-in’s letter, an anonymous call, an unattributed tip. The evaluator did not get to inspect the mechanism in those cases either. The operative question has never been “do I understand the source?” It is: what procedure extracts signal from a source I cannot explain, without my belief becoming a dependency of the analysis? The forecasting-tournament literature4,7 establishes that streams of probabilistic forecasts can be scored rigorously; what follows applies that insight where it has not been systematically applied — to the uninvited source with a dated archive and no acceptable pedigree.
2.5.2 The protocol
When a forecast stream arrives from a source whose mechanism cannot be evaluated:
Step 1 — Demand anteriority: a seal, not a story. The source’s own account of when a claim was made carries no evidential weight; a cryptographic or platform-issued timestamp carries all of it. The tooling is free and public: hash-commit the claim, anchor via OpenTimestamps, or rely on a server-issued timestamp the claimant cannot control (§2.1.2). A source unwilling to be sealed is requesting faith, and faith is not an admissible input to the procedure.
Step 2 — Freeze the claim verbatim. Not a paraphrase, not a summary, not the source’s later gloss — the exact words, hashed. Post-hoc reinterpretation is the central failure mode of anomalous-claim evaluation: an ambiguous sentence, re-read after events, can be made to “predict” nearly anything. A frozen verbatim text can be graded against outcomes by rules set before the outcome. (The record under study applies this discipline against its own interest: the published grading instruction scores the sealed pre-event claim text, never the descriptive title assembled afterward.) ( Figure 3.)
A vague claim covers nearly every outcome and cannot be wrong (≈ 0 bits of surprise-if-true); a specific claim names one outcome and dies if wrong. A forecast is worth what it excludes — the property Step 2’s verbatim freeze protects, and the reason post-hoc reinterpretation is the central failure mode of anomalous-claim evaluation. Authored vector plate (print edition), from the vector figure set of the companion monograph (CC-BY 4.0; Data availability).
Step 3 — Grade the stream, never the anecdote. One sealed hit is noise: indistinguishable from luck, base rates, or a rich prior. Demand the source’s whole record — every claim, sealed, with the misses kept in the denominator — and compute calibration over all of it. A source who offers highlights is offering nothing.
Step 4 — Let calibration, not theory, set the weight. One does not need to know why a source performs to know how it performs, and the second question is the only one a proper scoring rule answers. Weight the source’s next claim by its frozen record — exactly as one would weight a statistical model, a liaison service, or an analyst’s track record — and update as the record grows. Belief about mechanism never enters; the weight is an empirical quantity.
Step 5 — Rely on the frozen denominator as a filter. The protocol’s free byproduct is that it sorts uninformative sources automatically. Vagueness dies at Step 2 (an unfalsifiable claim cannot be graded); cherry-picking dies at Step 3 (the denominator is frozen); and a confidently wrong source grades itself into irrelevance within a modest number of entries, at essentially zero analytic cost to the evaluator. The protocol is not a courtesy extended to strange sources. It is a machine for sorting them, and it runs on arithmetic instead of belief.
2.5.3 Properties
Three properties are worth stating explicitly. Belief-independence: at no step does the evaluator assert, or need, any proposition about the source’s mechanism; the procedure’s outputs (a frozen record, a calibration profile, a weight) are invariant to the evaluator’s priors about the source. Symmetry: the protocol applies one standard to explained and unexplained sources alike — the same demand for anteriority, frozen text, and full denominators that a statistical model or a human forecasting team should meet. It neither privileges nor penalises a source for the acceptability of its mechanism. Low cost: every tool required (hashing, OpenTimestamps, a scoring script) is free and public, and the marginal cost of admitting one more sealed source to the procedure is near zero, because badly performing sources self-filter (Step 5).
The corpus examined in this paper is the protocol’s worked instance: a source with a maximally stigmatised mechanism, processed end-to-end through Steps 1–5, yielding the scored record of §3. The paper’s claim is not that the source is valid; it is that the procedure renders the question of validity empirically tractable without requiring anyone to believe anything.
Aggregate Brier = 0.0958 (log-loss ≈ 0.39). By desk: launch mission-assurance (n = 23) Brier 0.036 with zero launch misses; strategic warning/geopolitical (n = 69) Brier 0.116. On this high-confidence sample a base-rate baseline ties the aggregate Brier; we concede this explicitly — the information content is in specificity and lead time, not in the scalar score, which is why §3.3–§3.6 and the distribution axes matter more than the headline number. The eleven forecasts added in the version-9 cycle are a single sealed national-security session — the India national-security sitting of 16 August 2024 (§3.10) — graded 8 HIT, 2 NEAR, 1 MISS, so the aggregate rose by design as two NEARs and a retained miss entered the denominator alongside eight hits. ( Figure 4; Table 1.)
Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
The Russia–Ukraine theatre, the corpus this paper’s motivating case is drawn from, stands at 31 graded calls: 24 HIT, 1 NEAR, 3 PARTIAL, 3 MISS, average lead 238 days.13
Table 1 · CORPUS COMPOSITION AND FROZEN SCORES — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Table 1. Table 1 — the graded corpus by desk. Grades and probabilities are frozen in the Bitcoin-anchored grading ledger; the Brier is the mean of the frozen terms. Rendered live in the online edition (see pointer above).
Perfect calibration is the diagonal; the misses sit in their bands at full weight. Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
A base-rate or consensus-following strategy scores zero bits by construction, at any n. Dots are settled calls colored by what they earned — misses earn nothing. Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
The per-axis profile is the primary artifact; the composite is only a summary. Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
Failures and clustering published and re-runnable. Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
The non-zero reliability term is the misses scored at full confidence — surfaced, not smoothed. Authored vector plate (print edition) — the live, interactive, recomputing version is in the online edition.
The adverse outcomes fall in the lower-divergence, lower-consequence cells; the simultaneously high-decision-value, above-median-surprise calls resolved HIT. Error location, not an accuracy rate: the aggregate Brier of §3.1, with every miss at full weight, remains the accuracy claim. Authored vector plate (print edition), from the vector figure set of the companion monograph (CC-BY 4.0; Data availability).
The discipline cuts both ways: would-be hits are declined on principle alongside retained pre-corpus misses — the denominator is governed by protocol, not outcome. Authored vector plate (print edition), from the vector figure set of the companion monograph (CC-BY 4.0; Data availability).
Misses are kept on the chronology at full weight. No arc can be drawn after the fact: each seal date is fixed by the platform clock and third-party archives, and the record is Bitcoin-anchored against revision (§2.1.2).
Hairlines tether the most-cited laws to the sealed, graded cases whose citations name them — the constellation is wired to the record, not to theory.
Right: Murphy’s decomposition of the same Brier. Below: the surprise each call carried at seal. All values recomputed from the published calibration and graded corpus.
A Murphy decomposition grouped by exact pre-registered probability reconstructs the published score to the digit (§3.5), and its non-zero reliability term is precisely the misses scored at full confidence — surfaced, not smoothed. All 9 misses were sealed at 0.78; no call sealed at ≥ 0.90 has missed (33 such calls; the only two blemishes above that line are NEARs — IA-RU-026 at 0.90 and LA-018 at 0.925). ( Figure 5.)
The corpus per-call median Information Yield is on the order of seven bits (a 6.8-bit median — roughly a 1-in-110 surprise-if-true, capped at 1-in-a-million and never compounded across calls), concentrated in the against-consensus, long-lead, mechanism-named quadrant. A base-rate or consensus-following strategy scores zero bits by construction, at any n. The exact median recomputes live from the published artifact. ( Figures 6, 7.)
The current result is 51 of 68 strict-successful independent events; at the coin-flip floor the exact-binomial tail is p = 2.2 × 10 −5, and the break-even luck prior is ≈ 0.65 — the per-event luck probability a skeptic must grant to render the record unremarkable (i.e., to lift the tail above 0.05, one must assume every sealed event — a named-fleet storm, a Moscow terror window, a day-wide launch no-go — had at least a ~ 65% chance of landing by chance). Broken out by desk, the launch record is 10 of 12 strict events (p ≈ 0.019) and the Russia–Ukraine record 23 of 30 (p ≈ 0.0026); the India national-security sitting (§3.10) is a single event that graded as a strict failure — because one thread carrying two NEARs and a MISS cannot clear the all-HIT bar — a worked example of how the clustering is deliberately unkind to the forecaster. These figures should be cited with their caveats attached; the caveats are part of the claim. ( Figure 8.)
(1) Murphy decomposition.11 Brier = Reliability − Resolution + Uncertainty, grouping calls by exact sealed probability:
Two ceilings travel with this. First, resolution is modest and we do not inflate it: a coarse 12-value confidence ladder (§2.3.2) caps how finely outcomes can be separated, so a low resolution term is partly a property of the instrument, not only the forecaster. Second, a half-credit-outcome caveat: NEAR and PARTIAL are scored as fractional outcomes, so the Uncertainty term is the variance of a graded outcome variable rather than a pure 0/1 Bernoulli — the identity still closes, but “uncertainty” should be read accordingly. ( Figure 9.)
(2) Spiegelhalter’s Z.12 Z = −0.48. Because|Z| < 1.96, we cannot reject perfect calibration at the 0.05 level — but the honest reading is the opposite of a victory lap: at n = 92 this test has low power, so failing to reject is weak evidence, not proof of calibration. The slightly negative sign means the observed Brier sits marginally below its null expectation, well inside noise.
(3) Cox calibration regression.14 Fitting logit (outcome) = intercept + slope · logit(p) gives slope = 0.957 (ideal 1.0) and intercept = 0.054 (ideal 0). A slope near 1 says the forecasts are neither systematically over- nor under-dispersed; an intercept near 0 says there is no large systematic bias. Ceiling: the fit is on 92 points over a coarse predictor ladder, so the intervals are wide and the parameters should be read as reassuring-but-imprecise.
(4) Sharpness = 0.330. Sharpness measures how far the forecasts stray from the base rate. It is meaningful only beside calibration — one can be arbitrarily sharp and arbitrarily wrong — and is reported as a descriptive property of a confident record, explicitly not as a skill claim.
(5) Calibration-in-the-large = −0.018. The mean forecast minus the mean outcome: a slight under-confidence, small in magnitude and, given (2)‘s low power, not distinguishable from zero.
Taken together: the probabilities are close to calibrated with no large directional bias; the resolution is honestly modest and partly instrument-capped; and every reassuring number is bounded by the sample size and the coarse ladder. None of these is a discrimination claim against any other forecaster, and none escapes the standing ceiling that the outcomes are self-scored.
Effective domains ≈ 4.2 across 11 theatres. The graded calls span eleven nominal theatres (launch, Russia–Ukraine, US elections, India elections and national security, markets, and others). Nominal breadth over-counts when one theatre dominates, so we report the inverse-Herfindahl effective count, 1/Σsᵢ2: ≈ 4.2 effective domains. The gap between 11 nominal and ~ 4.2 effective is itself the honest disclosure — three theatres carry most of the weight — but ~4.2 is a generalist span, not a single-domain record.
100% per-call verifiability. Every one of the 92 graded calls publishes, in machine-readable form, its SHA-256 seal hash, its Merkle proof against the manifest root, at least one outcome source link, and a copy-pasteable verify command. Verifiability is complete by construction — and must be read for exactly what it is: verifiability is not accuracy. A fully verifiable record can still be a wrong one; the property proves only that no row asks the reader to take the author’s word.
An against-the-room trajectory. Classifying each call by whether it departed from the consensus base rate at seal time, the contrarian share rises from 17% in the first half of the record to 65% in the second half. This is descriptive, and the contrarian label is operator-assigned from published base rates (and thus contestable) — but the trend indicates the record moved toward, not away from, falsifiable against-consensus positions as it matured. (Exhibits 16, 19.)
EXHIBIT 16· THE LEAD-SPECIFICITY FRONTIER — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 16. Every graded call plotted on lead time versus specificity, misses included. Long lead is easy if vague; specificity is easy if late. Fifteen HITs occupy the ≥180-day, ≥80/100 region where neither shortcut exists. Rendered live in the online edition (see pointer above).
EXHIBIT 19· THE CAMPAIGN ARC — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 19. The 31-call Russia–Ukraine corpus read as one graded arc — each sealed call placed in the war’s chapters, leads of 4 to 792 days. Rendered live in the online edition (see pointer above).
Independent confirmation. A provenance property, reported separately and never entered into the Brier: six sealed calls were later confirmed not by the operator but by the subject of the forecast, an adversary, or an independent authority — Russia’s Ministry of Defence confirmed the strike on the landing ship Novocherkassk (IA-RU-023); Russia’s Foreign Intelligence Service publicly corroborated the sealed Ukraine-nuclear-intent call (IA-RU-020); the U.S. Embassy in Moscow issued its own public alert echoing the sealed Crocus window ~200 days after the seal (IA-RU-008); a public acknowledgement by Elon Musk bears on the Sevastopol plot advisory (IA-RU-006); the White House confirmed the Gershkovich prisoner-swap talks (IA-RU-017); and an AI system invited to out-forecast a sealed launch call publicly conceded the comparison (LA-014). Confirmation of an event is not adjudication of a grade — the grades remain self-assigned — but external parties, several of them hostile, independently affirmed the underlying events the sealed calls named. Reported as provenance, not as a skill statistic.
Corroboration density and value concentration. Every graded call carries public citations on its advisory page — a median of four independent sources, 386 in total — so the outcome side of each seal is a checkable dossier rather than a bare assertion (a documentation axis, not a skill axis). Cross-tabulating decision-value (SITA) against surprise (Information Yield), the calls that are simultaneously high-decision-value and above-median-surprise all resolved HIT, while the record’s adverse outcomes fall in the lower-stakes, lower-surprise cells. This is error location, not an accuracy rate: the aggregate Brier of §3.1, with every miss at full weight, remains the accuracy claim; the cross-tabulation adds that the record’s errors concentrate where the decision weight is lightest. ( Figure 10.)
Method breadth beyond the scored theatres. The effective-domain count is a floor on the method’s applied range, because the discipline declines to score some theatres it nonetheless reads: an 11 November 2023 sealed seven-part reading of the Israel–Hamas–Palestine conflict contained an institutional call that materialised as UN Security Council Resolution 2728 (25 March 2024) — a mechanical HIT held ungraded on ethics (§6.8).
The distribution axis. Seal integrity proves the text did not change; grading proves the misses stayed; distribution proves there was no retreat path. The Russia record accumulated 50.6 million impressions; across all four owned channels the documented public reach is 238.2 million impressions, of which ≈ 124.3 million lands on the graded-advisory posts specifically (the exposure metric, distinct from total reach). Every figure is a floor, not a ceiling: the platform reports impression_count only from ~December 2022 onward, so all earlier posts count zero. The single sealed warning that preceded the March 2024 Crocus City Hall attack accumulated 653,566 impressions before the event, and one sealed thread opener reached 5.93 million. Organic reach on these accounts was effectively zero — self-reposts register zero impressions in the same dataset — so distribution was paid, worldwide, and pre-event. This is methodologically relevant, not promotional: a dated forecast shown to millions before its event cannot be quietly deleted or walked back, and hostile audiences preserve independent copies. The sealed Russia–Ukraine subcorpus was additionally released as a self-produced documentary film in 42 simultaneous language versions (jyotishintelligence.com/film), noted here strictly as a distribution-breadth and anti-revision property: simultaneous mass release across dozens of language editions forecloses quiet post-hoc revision of the narrative record, complementing at the social layer the tamper-evidence the cryptographic seals provide at the artifact layer. Reach proves exposure and falsifiability, never accuracy; the ceiling travels with the claim.
The delivery record. Beyond publication, the recovery documents dated public delivery of the forecasts to named institutional recipients, each its own timestamped, linkable post: the full sealed thread posted to the public account of the Deputy Chairman of Russia’s Security Council on 2 January 2024 and again on 17 March 2024 (five days before the Crocus attack); to Russia’s Deputy Permanent Representative to the United Nations (September 2023); a narrowed 2026 nuclear-test window into the public threads of Russia’s Ministry of Foreign Affairs and state media (February 2024); and assessed locations to a U.S. Army counterintelligence account (March 2026). The record of who was publicly addressed, and when, is as verifiable as the forecasts. Consistent with §1.1, no receipt, readership, or action by any recipient is claimed. (Exhibits 15, 20.)
EXHIBIT 15· COMPARATOR ABSENCE — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 15. For most of the warning record no official public warning existed at any lead; where dated comparators exist, the seal preceded them — with the operationally superior comparator conceded first. Rendered live in the online edition (see pointer above).
EXHIBIT 20· TWO WARNINGS, ONE PROTOCOL — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 20. The same seal protocol across jurisdictions and event classes; graded strictly where gradeable, held ungraded where grading would claim credit the record refuses (prevention-paradox class). Rendered live in the online edition (see pointer above).
The materialisation log. Running alongside the delivery record is a chain-of-custody log: across the Russia–Ukraine corpus the operator posted, in real time as each call resolved, dated public comments tying the sealed forecast to the reported outcome at the moment it occurred. Each entry is an independently timestamped public post, so the resolution trail — not only the seal — is third-party reconstructable from platform-native metadata. The log is descriptive provenance and enters no score.
The paper’s motivating sentence did not arrive alone, and the manner of its arrival is itself evidence, so this subsection works the sitting that produced it end-to-end.
On 4 September 2023, a single session emitted one numbered 59-post public thread, every post server-timestamped within the same hour (the thread’s status IDs decode, by the arithmetic of §2.1.2, to 4 September 2023, ~18:09 UTC onward). Fourteen distinct graded futures carry that one day’s seal. Three are marquee entries in their own right:
• Post 37 — the Crocus window (IA-RU-008, HIT, ~200-day lead). The verbatim sentence of §1.1: city, target class, two-month window, stated intent.
• Posts 5–8 — F-16 basing (IA-RU-010, HIT, ~301-day lead). The thread named Starokostiantyniv — a specific airbase in Khmelnytskyi Oblast, one of at least five candidate fields — as the operating base for Ukraine’s not-yet-delivered F-16 s, and framed the aircraft as a standoff asset rather than a war-winner. The aircraft arrived in mid-2024; the base was corroborated by repeated strikes against it and by open-source imagery, roughly 301 days after the seal.
• The Black Sea storm (IA-RU-006, HIT, 83-day lead). The same sitting sealed a warning of a destructive wave event against the Crimean coast and the Black Sea Fleet’s basing. Eighty-three days later, on 26–27 November 2023, the widely reported “storm of the century” struck the Black Sea and the Crimean coast with exactly that geography. For calibration: modern numerical weather prediction, the most mathematically mature forecasting enterprise available, holds deterministic skill to roughly ten days15; no meteorological method forecasts a specific storm at 83 days. Whatever produced that entry was not weather forecasting — which is precisely why it is scored, like everything else, only against its sealed text.
The same sitting also produced, at post 53 of 59, a call that Roscosmos would launch a Mars rover. That call failed and is on the ledger as IA-RU-032, graded MISS. So is a second miss from the same sitting: a Duma-succession call naming specific figures that did not materialise (IA-RU-028, MISS). The structural point deserves emphasis, because it is the record’s sharpest answer to the cherry-picking hypothesis: the sitting was frozen whole — 59 numbered posts, crown jewels and embarrassments under the same hash discipline, sealed in the same minutes — and two of the Russia theatre’s three misses live in the same sealed artifact as its most-cited hit. That is not humility performed in prose; it is humility a reviewer can recompute.
Across all six multi-call seal dates, 44 graded calls resolve to 36 hits: the sitting is a high-yield unit, but it is not uniformly successful, and its failures are retained beside its successes at full weight — one sitting, both outcomes, no halo. The same discipline applies to the two India sittings: the 10 August 2023 reading, whose 9 graded calls include eight hits and one clean MISS (the economy call, IA-IN24–009) sealed in the same minute, and the 16 August 2024 national-security sitting (§3.10), the sixth and most cross-domain, whose 11 graded calls run 8 HIT/2 NEAR/1 MISS. (Exhibit 14.)
EXHIBIT 14· THE SITTINGS — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 14. Complete enumeration of multi-call seal dates — six — 44 graded calls of which 36 are HITs, resolving 4 to 752 days later. The 4 September 2023 sitting alone sealed 14 distinct futures, its two misses (the Duma succession plan and the Roscosmos Mars rover) kept on the ledger at full weight. Rendered live in the online edition (see pointer above).
Eleven graded additions
The full-corpus recovery surfaced eleven previously un-catalogued sealed calls that met the grading criteria (a twelfth, an F-16 employment-pattern call, was withdrawn to avoid double-counting a sealed sentence already scored under the F-16-basing advisory). They were added at full weight, misses included:
The Kursk entry deserves one methodological note, because it demonstrates that lead time is not the record’s only axis: sealed 7 March 2025 (status ID 1898052798266458420, snowflake-decodable to the day), it placed a dated deadline — collapse by 17 March — on a contested front ten days out; Sudzha, the salient’s anchor town, fell between 12 and 16 March 2025. Long lead is one form of specificity; a short-fuse dated deadline is another; the ledger carries both.
A pre-stated falsification condition fired
The record published, in advance, the conditions that would falsify it — among them, in its published wording, “a MISS on either buyer desk” (i.e., on either applied desk: launch mission-assurance or strategic warning). Grading IA-RU-022 (the 9 May 2025 ceasefire window) met that condition on the strategic-warning desk. The standard responses — reword the condition, re-scope the desk, or let the claim lapse — are all foreclosed by the protocol. The condition’s wording is preserved verbatim in the published artifact; the miss entered the Brier at full weight (its 0.6084 term is tied for the heaviest in the corpus); and a previously true “zero misses” claim was retired across every surface. We report this first, not last, because it is a methodological strength: a record whose falsification conditions cannot fire was never falsifiable, and a fired-and-retained condition is the demonstration that the file-drawer and post-hoc-revision problems of §1.3 are actually foreclosed, not merely disclaimed. (Exhibit 21.)
EXHIBIT 21· ANATOMY OF THE MISSES — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 21. All nine misses, at full ledger weight, sealed at a uniform 0.78 across hits and misses: four electoral, three on the Russia–Ukraine warning desk (IA-RU-022, −028, −032), and two on the India desk (IA-IN24–009, IA-IN24-NS09). The record has never missed at ≥0.90. Kill-conditions are pre-stated and armed — and one has fired: the buyer-desk-miss condition (IA-RU-022), owned in the open rather than reworded. Rendered live in the online edition (see pointer above).
The corpus funnel — an auditable selection result
The most common objection to any curated forecast record is “you only show a subset — where are the ones you got wrong, or the ones you never sealed?” The recovery answers it directly for the largest theatre. The classes below sum exactly to the 1,031 original posts (374 retweets excluded upstream); of these, 612 are numbered thread sub-forecasts that cluster into the 31 graded advisories, and the remainder carry no dated, falsifiable, standalone claim:
The reduction from a four-figure posting record to 31 scored advisories is dominated by thread-clustering and non-forecast content, and the only distinct resolved event outside the 31 is excluded on ethics, not hidden as a win. The funnel converts “you’re cherry-picking” from an assertion into a checkable table.
The India national-security sitting (16 August 2024)
The record’s clearest single test of compression — how many independent, dated, falsifiable calls one sealed reading can carry, and how they grade when scored together — is a single sealed thread issued on 16 August 2024. It named a 24-day national-security window for India (27 August–18 September 2024) and partitioned it into eleven separate falsifiable sub-forecasts, each sealed verbatim in the same thread and each graded afterward against the public record as its own advisory (IA-IN24-NS01 … 11). The eleven span deliberately different domains, so no single world-state drives them: border, cross-border terror, festival-period public safety, military posture, protectee security, industrial/environmental disaster, civil unrest, government stability, financial markets, public health, and the space programme. This section is written in a deliberately apolitical register: only the sealed class of each call is graded, no partisan wording is relied upon, and the two sealed sub-phrases that carry sensitive connotations are neither quoted nor interpreted here.
Graded strictly, the sitting resolved 8 HIT, 2 NEAR, 1 MISS:
The decision-value frame. The value of a sealed warning is not its scalar score but the decision it would have informed, dated and specific, before the event: a 5-day lead on an industrial/chemical disaster spate relevant to plant-safety regulation; a “targets beyond Jammu & Kashmir” terror flag relevant to rail and transport security 23 days out; a festival-window public-safety flag 26 days out; a protective-coverage flag around a named overseas engagement 37 days out; a public-health surge flag 32 days out. Each is falsifiable and each carried an actionable lead — which is precisely why the two NEARs and the one MISS are surfaced first, not last. (No claim is made that any of these informed any decision; §1.1.)
The two NEARs, kept as NEARs. NS05 is graded NEAR, not HIT, deliberately: the sealed heightened-risk period materialised as a credible, documented public threat and a raised protective posture — but no peak event occurred, and the record will not inflate a materialised threat into the sealed peak event that did not happen. (The call concerns protective coverage; it is never a death forecast, and is not read as one here.) NS07 is graded NEAR because the sealed issue-set was demonstrably live in the window while the specific claimed mechanism — a protest wave forcing a concrete legislative change — did not discretely materialise: themes matched, the causal event did not.
The MISS, surfaced not buried. NS09 called a failure within the financial system and consequent market instability. The opposite happened: through and just after the window, Indian equities climbed to records (the Sensex first crossing 85,000 on 24 September 2024). It is an inverted call — the market was at peak strength, not distress — and it is retained on the ledger at full weight, sealed at 0.78 like every other call in the thread. Under the squared-error rule (§2.3.6) the sitting’s Brier is (8·0.0484 + 2·0.0784 + 1·0.6084)/11 ≈ 0.105 — visibly worse than the corpus aggregate, and worse because of one call: the single miss contributes more to the sitting’s total squared error (0.6084) than the other ten calls combined (0.544). A scoring rule under which one honest miss outweighs ten hits cannot be farmed by volume; the only way to move it is to be right at the confidence actually claimed.
Other illustrative calls
Crocus City Hall (IA-RU-008), HIT, ~200-day lead. Sealed 4 September 2023 (§1.1, §3.8). The sealed text named city, target class, window, and psychological intent; it named no venue and no perpetrator, and only what was sealed is graded. The paper makes no attribution claim. The warning was distributed and publicly delivered before the event (§3.7).
New Glenn NG-3 (LA-022), HIT, 1-day lead. Sealed 18 April 2026, the day before the window: flagged the second launch window as high risk of catastrophic outcome and named payload loss; the booster recovered but the payload was lost and the vehicle grounded.
Kursk front collapse (IA-RU-021), HIT, 10-day dated deadline. §3.9.1.
Tactical-nuclear consideration window (IA-RU-009), HIT. Sealed a consider-then-rule-out dynamic in a 15–31 May 2024 window; graded on the sealed scenario (considered, not used), not on a detonation.
Maharashtra 2024 (IA-MH24–001), MISS. A 288-seat forecast sealed nine days before the vote; the headline alliance call was wrong and is retained at full weight. Context, not mitigation: the entire 16-poll professional field also missed, above every exit poll’s ceiling.
Doha death-row read (IA-INGEO-001), HIT, ~109-day lead — the record’s consular-intelligence entry. Sealed 26 October 2023, four days after a Qatari court sentenced eight former Indian Navy veterans to death, the read called a favourable resolution reachable within a window closing end-May 2024; the sentences were commuted and the men released (February 2024), inside the window. Only the embedded chart-based forecast is graded, not the diplomacy: the outcome is properly credited to Indian diplomatic effort and the Qatari courts, and the paper makes no causal claim. Its Information Yield is modest (≈ 2.6 bits) and it is not classified against-consensus in the dataset; what it adds is breadth.
India economy slowdown (IA-IN24–009), MISS, ~295-day lead. Sealed 10 August 2023 in the same 55-post reading that produced eight India-2024 hits, at 0.78: a macro-direction claim that the economy would “take a thorough beating and slow down.” The macro-direction was wrong; it is graded a clean MISS and retained at full weight beside its eight thread-siblings — the sitting is high-yield, not infallible.
A boundary stated live. The record’s edges are as informative as its centre. On 24 June 2023, during the Wagner Group mutiny, the operator posted — while the outcome was undecided — “I did not predict a coup because there will be none,” and later that day, hours before resolution, “This rebellion will fail” (status IDs 1672397949450739712 and 1672591213336346624, both decodable to the hour by the arithmetic of §2.1.2). Both were correct. But twice in those same hours the operator stated, in public, the thing a fabricator never states: “I did not predict the mutiny.” The sealed record does hold a March-sealed escalation window for late June, catalogued afterward under the mutiny’s name, but it contains no prediction of a mutiny as an event, and that boundary was drawn by the operator, twice, in public, while the outcome was still undecided. A record that claims everything is a record that means nothing; the boundaries this corpus draws on itself are part of its evidence.
(Exhibits 7–13 and 22 present the launch-desk decompositions preserved from the prior version: signal-detection structure with zero false all-clears, the mechanism ledger of named anomaly classes, the dose-response ladder, provider neutrality across five independent organizations in two countries, the Axiom-4 repeated-attempt series, and the Falcon 9 same-vehicle isolation.)
EXHIBIT 7· SIGNAL-DETECTION DECOMPOSITION — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 7. Every launch call by day-level operative direction. Zero false all-clears; zero missed catastrophes; the only errors are two self-penalized severity overcalls. Descriptive of the graded record, not a validated-detector claim. Rendered live in the online edition (see pointer above).
EXHIBIT 8· THE MECHANISM LEDGER — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 8. All 23 launch rows: the anomaly class named in the sealed text against the realized outcome. Zero bare scrub calls — every risk-flagged row names a subsystem or anomaly class. Rendered live in the online edition (see pointer above).
EXHIBIT 9· THE DOSE-RESPONSE LADDER — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 9. The full 23-row intensity ladder, sealed language against outcome severity, with the two gradient-breakers (LA-018, LA-012) shown in place. No coefficient is published: the ordering is authored, so the ladder is shown, not scored. Rendered live in the online edition (see pointer above).
EXHIBIT 10· PROVIDER NEUTRALITY — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 10. The launch record across five independent, competing organizations in two countries. Insider-family explanations are provider-specific; this table requires all five at once. Rendered live in the online edition (see pointer above).
EXHIBIT 11 · AXIOM-4 — THE CLEAN EXPERIMENT — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 11. Same booster, capsule, and crew across four dated attempts; hardware improved monotonically as faults were repaired; the sealed reads varied against that gradient and graded 4/4. The reads did not track machine state. Rendered live in the online edition (see pointer above).
EXHIBIT 12· FALCON 9 — THE GOD-VEHICLE ISOLATION — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 12. The most-debugged orbital vehicle of its era as the cleanest isolation of day-level variance: six sealed reads, six HITs, four different verdicts on an identical stack within fifteen days. Granularity, not channel. Rendered live in the online edition (see pointer above).
EXHIBIT 13· HARDWARE DECOUPLING — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 13. Sealed read versus machine state across the corpus — the reads decouple from hardware condition, the signature a day-level signal leaves and a hardware-inference strategy cannot. Rendered live in the online edition (see pointer above).
EXHIBIT 22· SUPPLEMENTARY LAUNCH-DESK EXHIBITS — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 22. Exhibit S1 — sub-claim self-downgrades inside winning composites, operational tempo, full failure-taxonomy coverage, named financial exposure, the intervention arm, and the cheap-scrub rebuttal counted. Rendered live in the online edition (see pointer above).
This section states what a sealed, scored, belief-independent forecast stream offers two analytic settings. Both are offered as tradecraft, with no counterfactual claims attached: nothing here asserts that any institution should have used, or would have benefited from, this record. And one ceiling is stated as permanent rather than modest: nothing in this record is, or could responsibly be, a go/no-go decision input for any operational decision. The record claims analytic grain only (§4.1), and the paper treats that as a structural property of this class of evidence, not a humility to be negotiated away.
The honest objection to the motivating case deserves a subsection rather than a footnote: what is an analyst supposed to do with “public gatherings, Moscow, March–April”? It cues no raid; it hardens no venue; it names no cell, no weapon, no date. A two-month city-wide window is not actionable in the operational sense, and if this paper implied otherwise it would deserve rejection.
So we state it flatly: the record never claims operational grain. The sealed sentence is analytic grain — a city, a target class, a window, a mechanism of intent. What analytic grain can do is re-weight attention: raise a standing question’s priority, shade a collection emphasis, lower the threshold at which fragmentary reporting on a matching pattern receives a second read. Analysts re-weight on soft inputs constantly — a liaison service’s mood, a defector’s aside, an editorial in a controlled press. The operative question is never whether an input is soft; it is whether the input’s track record earns it weight — a question that can only be answered if somebody kept score. The source-agnostic protocol of §2.5 is the keeping of that score, generalised: it gives an evaluator a procedure for the unsolicited input that requires no belief, costs almost nothing, and automatically disposes of the sources that would waste attention.
The instrument that makes Step 3 (“grade the stream”) practical at corpus scale is worth one note, because grading a stream is easy to prescribe and tedious to do. The record under study publishes its stream not only as files but as an interactive audit surface: every sealed advisory rendered with its seal provenance and Brier contribution attached, a temporal replay that runs the record forward through time (each call appearing on the date it was sealed, before its event resolved, so a reviewer can watch the stream accumulate as a contemporaneous observer would have), and deep-linkable view states so one reviewer can send another the exact filtered, time-scrubbed state under discussion. The surface is not the evidence; it is the reading room, with the evidence hashed to the walls. Auditing the Russia–Ukraine theatre against events takes an afternoon, from the seals, with the author absent from the trust chain.
War games inherit the warning problem one level up: a game can only explore scenarios someone imagined, so the scenario list is bounded by the planners’ priors, and the failure-of-imagination problem is not merely present in gaming — it is institutionalised by it. A sealed, scored, against-consensus forecast record bears on that problem in three ways, in increasing order of generality.
Scenario selection. A gaming cell must decide which windows merit convening a game at all and which failure mechanisms to inject. A graded record of against-consensus calls is a natural red-cell seed for that decision, with a property brainstormed injects never have: accountability. An inject drawn from a sealed record arrives with its author’s public calibration attached; the cell can weight the scenario by the record rather than by the seniority or eloquence of whoever proposed it. The Kursk entry (§3.9.1) illustrates the shape of the contribution: a sealed, dated read of a vulnerability window — a front’s collapse, with a deadline no public forecast had placed — scored after the fact against reality. The claim is only that reads of this shape exist, dated, in a graded record, and that scenario menus are where such reads belong. Windows, always — never targets: a graded record contributes when a system is ripe to break, which is different in kind from any question of how to break it, and this paper contributes nothing to the latter.
Parameterization — the escalation matrix. A game’s verdict is only as good as the values inside it: the event likelihoods, the timing windows, the assumed actor responses. Most games run on planner guesses and consensus priors, and this is a large part of why game outputs diverge from what reality later does — the machinery is sound and the inputs are ungraded. The sharpest form of the problem is the escalation matrix: the lattice that maps, for every move a game examines, how the adversary answers — militarily, economically, diplomatically, in cyber, in the information space. Every branch of that lattice carries an assumed likelihood, a threshold, and a clock, and the matrix is only as sound as those values, because which branch an adversary takes is not a setting; it is a forecast. An input source that carries a public calibration record — a Brier of 0.0958 over 92 frozen entries, misses kept, luck-tested against clustered-event nulls — is a different class of object from an unscored estimate, whatever its origin, because the one thing known about it is exactly how wrong it has been. Nowhere does this matter more than in nuclear gaming, for a reason worth stating plainly: it is the one domain with, mercifully, almost no empirical events to calibrate against. A conventional game can at least be graded against the last war; a nuclear game runs on assumption all the way down, reality’s adjudication never arrives, and the cost of a mis-parameterized branch is not a bad afternoon in the game cell but a decision made years later on the game’s remembered verdict. In precisely that domain — maximum stakes, minimum data — input provenance and sealed-assumption scoring are not refinements; they are the only calibration such a game can ever have. The payoff runs in both directions: populate a matrix’s branches with scored, calibration-tested values instead of unscored estimates, and the game’s realism stops being a quality asserted by its designers and becomes a property inherited from its inputs.
Adjudication tradecraft — sealing the game itself. The third application is free, requires no unconventional source, and may be worth more than the other two combined: apply the sealing protocol to the games themselves. Before events, seal the game’s key assumptions, its parameter choices, and its expected outcomes — hash, anchor, timestamp, exactly the mechanics of §2.2. After events, grade them under a rubric fixed at seal time. Within a few cycles, a gaming cell owns something almost no institution possesses: its own frozen calibration record — a quantitative answer to “how often do our games’ assumptions survive contact with reality?” — with the same anti-self-deception property the anchor gave this record’s misses: nobody can quietly forget the assumptions that failed. This adopts the entire tradecraft of this paper while adopting nothing of its source, and it costs a hash and a timestamp.
The institutional logic of beginning in the gaming cell is worth making explicit: a war game is an exploratory sandbox, not a formal estimate. Feeding a fully sealed, fully scored, unconventional input into a game commits no one to anything — no estimate is signed, no decision leans on it, and the input’s record is graded alongside everyone else’s when reality arrives. An exploratory venue whose purpose is to consider what the consensus has not is precisely where a belief-independent scoring protocol gets its first fair test.
The discipline underneath everything in this paper can be stated in one sentence: never grade the anecdote; grade the stream. The single striking forecast — a 200-day terror window, an 83-day storm with named geography — is uninterpretable in isolation: indistinguishable from luck, base rates, or a rich prior, and no amount of narrative can change that. A frozen stream with its denominator intact is a measurable object. This is the forecasting-tournament insight4,7 applied where it has not been systematically applied — to the unsolicited source, the unexplainable input, the record that arrives without a pedigree — and it reverses the usual burden: the question stops being “can this source be believed?” (unanswerable, and the wrong question) and becomes “what does this source’s frozen record measure?” (answerable, by arithmetic).
The corollary governs how this paper should be cited. The Crocus case is the record’s most legible artifact, and it is the wrong unit of citation. The correct units are the corpus statistics of §3 with their caveats attached — the base-rate tie conceded, the luck test’s hostile construction stated, the self-grading ceiling in view. The record’s evidentiary weight, whatever it is, lives at the level of the stream, and only there.
The strongest single objection to the motivating case deserves its sharpest statement. Wartime Moscow; elevated terror risk by construction; a two-month window; “possible” as the operative verb — is the Crocus entry anything more than an informed reading of ambient conditions? For the single call, standing alone, the objection has real force, and we decline to answer it at the level of the single call. The answer lives at the level of the stream.
A base-rate follower — a forecaster mechanically restating ambient risk — produces a specific, reproducible statistical signature: hedged windows, no mechanisms, leads clustered near zero, and zero Information Yield by construction (§2.3.3). Run that strategy against this ledger and it can approach exactly one number: the aggregate Brier — the record’s own documentation concedes as much, and this paper concedes it in the abstract. What the strategy cannot produce, by construction, is what the frozen record shows elsewhere: an 83-day storm call with named geography; a ten-day dated deadline on a contested front that resolved inside its window; a specific airbase named ~301 days before the aircraft operated from it, selected from a field of candidates; and intent language (“sow panic amongst a public well-insulated”) fixed in the sealed text before the event. Base rates do not produce named mechanisms, named geography, or dated deadlines. The ledger — with Information Yield, the lead-specificity frontier (Exhibit 16), and the luck test’s strict event scoring — is the instrument that measures the difference without requiring anyone to name its cause.
This is also why the paper reports the base-rate tie so prominently rather than burying it: the scalar Brier is the one axis on which the record and a base-rate strategy are indistinguishable at this sample size. Every other reported axis — specificity, lead, yield, the strict-event tail — exists precisely to interrogate the axes on which they are not.
What it shows: that a complete, dated, verbatim-frozen, publicly distributed forecast stream can be maintained under a strictly proper scoring rule with misses retained and falsification conditions armed; that such a stream can be audited end-to-end by a hostile reviewer at near-zero cost, offline, with the author absent from the trust chain; and that this particular stream’s frozen record contains a combination — long-lead specific hits, a strict-event tail of p = 2.2 × 10−5 under hostile clustering, honest misses at full weight — that a base-rate strategy cannot reproduce on the non-scalar axes. One sentence of economic context bounds the practical stakes of the launch subset: the 23 launch-day forecasts covered missions carrying approximately $6.2 billion of publicly documented cost (per NASA OIG, GAO, and SEC records), of which exactly one asset was commercially insured — context for the practical asymmetry between the cost of evaluating an out-of-model signal and the cost of the days it addresses, and not a prevention or counterfactual claim of any kind (§1.1, §4).
What it does not show: that the generative method is valid (no mechanism is offered and none is claimed); that the operator’s probabilities are well-calibrated in a statistically demonstrated sense (the sample is small, the suite under-powered, the base-rate tie conceded); that the record would perform prospectively under external adjudication (that test is defined but not yet run; §6.7); or that any warning here was, or should have been, acted upon by anyone (§1.1). The record renders a stigmatised question empirically tractable. It does not settle it.
This section is deliberately the paper’s most load-bearing. Each limitation is stated at full strength, with its mitigation where one exists and without one where none does.
Probabilities and grades are operator-assigned. No independent party has adjudicated the record. The mitigations are structural, not rhetorical: the rubric is frozen at seal time; the inputs (verbatim claim, outcome, sealed probability, verdict, base rate, specificity vectors) are published; and a public regrade kit plus the CC-BY corpus let any reader substitute their own probabilities and verdicts and recompute the Brier. This converts “trust the grader” into “grade it yourself” — independent regrading is not merely possible but invited; a hostile grader can run the exercise and publish a different mean — but it does not substitute for third-party adjudication, which remains outstanding. “Audited-but-not-peer-adjudicated” is the accurate description of the record’s current status.
EXHIBIT 23· THE ANTI-BARNUM CURVE — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 23. Named falsifiable elements per sealed call across the corpus, rising roughly monotonically from one to twelve while accuracy held — the inverse of the documented charlatan trajectory of retreat into vagueness under scrutiny. Rendered live in the online edition (see pointer above).
n = 92 is small and weighted toward high-confidence calls; a base-rate baseline ties the aggregate Brier, and the calibration-integrity suite (§3.5) is explicitly under-powered at this size (Spiegelhalter|Z| < 1.96 reflects low power, not proven calibration). The against-consensus subset (Information Yield >0) and the strict-event luck test are the more informative reads, and both are reported with their caveats.
Which matters get sealed is chosen by the operator, and the seal-time selection process is not externally auditable in principle: a reader can verify that every sealed call is anterior and retained, but cannot verify what was considered and not sealed. The corpus funnel (§3.9.3) resolves the selection question for what was posted; it cannot resolve it for what was never posted. This is the single hardest limitation and it is not resolved by the recovery. Only a denominator-fixed prospective season with external adjudication retires it.
The §3.7 reach figures establish that the forecasts could not be quietly withdrawn; they say nothing about whether the forecasts were correct. Paid distribution is evidence of commitment and falsifiability, not of skill. The axes are separated deliberately, and the separation should be preserved in citation.
The disclosed method is Vedic jyotish. The paper makes no claim about the validity of that mechanism, offers no causal account, and treats the generator as a black box (§2.5.1). The contribution is entirely in the protocol and the graded record; a reader who rejects the method entirely can still evaluate the anteriority, the retention of misses, and the reproducibility, none of which depend on the method being valid.
The recovery (§2.4.2) confirms when posts existed via platform-native timestamps (snowflake IDs and archive capture), independent of when any descriptive page about them was published. A reviewer should note that narrative framing on the live site postdates the events; the sealed artifacts do not, and only the latter carry evidentiary weight. For events that resolved before the full-corpus manifest was Bitcoin-anchored, anteriority rests on the platform clock, pre-event distribution, and third-party archives rather than on the blockchain anchor (§2.1.2); the two-clock seam is disclosed rather than papered over.
This is a retrospective, self-adjudicated record. A standing public board of sealed-OPEN forward calls exists, and the protocol’s design goal is a prospective, externally adjudicated, denominator-fixed test; but that test is not what this paper reports. The record runs forward under the same frozen rules — the next entries will resolve as hits or as misses, and the record’s integrity consists in the operator being unable to do anything about it either way — but forward performance is a question for a future paper, not a claim of this one.
Selection cuts both ways: some sealed material is deliberately held out of the graded denominator, with the exclusion reason published, precisely because scoring it would be dishonest — not because it would score badly.
• Named-individual personal-safety forecasts (e.g., a 2025 head-of-state safety read; the “distinct resolved event not in the 31” of §3.9.3). Excluded on two grounds: an unfalsifiable prevention paradox (if a warned-of harm does not occur, one cannot distinguish a wrong forecast from a prevented one) and ethics (the record does not publish scored death-forecasts about named living individuals).
• Live-catastrophe reads that are tracking. Even where a sealed read appears to be matching events, an ongoing humanitarian catastrophe is kept out of the scored corpus on ethics: scoring live suffering as a forecasting win is a line the protocol will not cross, regardless of accuracy.
• A would-be HIT, declined on ethics. An 11 November 2023 sealed seven-part reading on the Israel–Hamas–Palestine conflict named a January–March 2024 institutional-governance development that materialised as UN Security Council Resolution 2728 (25 March 2024), clearing the mechanical bar for a graded HIT. It is nonetheless ungraded by choice, on the same live-catastrophe ethics.
• Advisory-register posts (the single off-protocol “advice/remedy” line in the funnel). Counsel is not a dated falsifiable claim; scoring it would inflate the denominator with non-forecasts.
That these exclusions remove apparently favourable and high-salience material — rather than unfavourable material — is the strength of the argument: a curation rule that discards potential wins on principle is evidence the denominator is governed by a protocol, not by outcome. ( Figure 11.)
Every headline number recomputes from public artifacts with no author in the trust chain. The fastest path is the two commands of §2.2; each step below removes one class of trust. (Exhibit 24.)
EXHIBIT 24· THE FEYNMAN LOOP — Presented live: the interactive, recomputing rendering of this exhibit is in the online edition at jyotishintelligence.com/essays/sealed-before-the-event-the-paper — every value recomputes from the published data artifacts named inline.
Exhibit 24. The record’s operating cycle mapped onto guess → compute consequences → compare with experiment — with the sealing protocol enforcing that the guess cannot be revised after the comparison. Rendered live in the online edition (see pointer above).
(a) Reproduce the whole paper — two commands. From the public repository github.com/vijayjyotish/verify-jyotint (both scripts zero-dependency, Node ≥18, offline-capable): verify-jyotint.mjs proves the record intact — integrity, on the anchor clock; the seal dates themselves are established on the platform clock by step (d) below; reproduce-paper.mjs proves the analysis an honest projection of the record. Together they regenerate the paper. A narrated walkthrough of every check, each with its exact expected output, is published alongside the corpus.
(b) Re-derive every metric from the flat table. Pull the one-row-per-call CSV and recompute:
import pandas as pd
df = pd.read_csv(“ https://jyotishintelligence.com/dataset/jyotint-analyst-table.csv “)
brier = df.brier_term.mean() # ≈ 0.0958.
(c) Re-run the clustered luck test. Take the published clusters, disagree with any clustering choice, re-partition, and re-run — the exact-binomial tail recomputes under any partition the reviewer prefers.
(d) Decode any advisory’s seal date from its post ID, by pure arithmetic — no API call, no trust in the author: timestamp_ms = (object_id > > 22) + 1288834974657 (e.g., IA-INGEO-001, object_id 1717617022161625165 → 26 October 2023, before the outcome; IA-RU-008, 1698760248033476914 → 4 September 2023).
(e) Recompute the calibration-integrity profile from the published measurement artifact (n, Brier, the Murphy triple, Spiegelhalter Z, Cox slope/intercept, sharpness, calibration-in-the-large) and confirm the Murphy identity closes: Reliability − Resolution + Uncertainty = the published Brier.
(f ) Inspect the pre-stated falsification conditions and the one that fired. Confirm that IA-RU-022 met the published buyer-desk-miss condition and is logged at full weight, its wording un-reworded (§3.9.2).
This paper set out to make one class of question empirically tractable: what to do, procedurally, with a forecast source that appears to perform but whose mechanism no evaluator can accept on its face. The answer defended here has two halves. The first is infrastructural: a sealing protocol — verbatim text, canonical hash, blockchain anchor, frozen rubric, retained misses, recoverable distribution — that converts a public forecast stream from a narrative into a measurable object, checkable end-to-end by a hostile reviewer at near-zero cost. The second is procedural: a five-step source-agnostic scoring protocol under which belief in the source never becomes a dependency of the analysis — anteriority demanded, claims frozen verbatim, streams graded rather than anecdotes, weights set by calibration, and uninformative sources filtered automatically by their own frozen denominators.
The record examined is the protocol’s worked instance, and its numbers are reported with their ceilings welded on: 92 graded calls, 73 hits, 9 misses at full weight, an aggregate Brier of 0.0958 that a base-rate baseline ties, a hostile clustered luck test at p = 2.2 × 10−5 with its break-even prior published, a five-way calibration suite that is reassuring but under-powered, and a fired falsification condition kept in its original wording. The generative method is disclosed and black-boxed; its validity is not claimed and cannot be established by this design. What the design establishes is the procedure by which that question — and the analogous question for any unexplainable source, in any evaluative setting from a warning desk to a game cell to a referee’s desk — can be answered over time by arithmetic rather than by authority: seal first, grade everything, keep the misses, and let the record say what it says.
The record runs forward under the same frozen rules. Whatever the next entries resolve to, the arithmetic will be waiting where any reader left it — in the capsule, in the ledger, in the block — timestamped, unalterable, and indifferent to opinion, including the author’s.
Portions of this manuscript were drafted, structured, and edited using a generative AI assistant — Claude (Anthropic; model Claude Fable 5, model ID claude-fable-5), accessed via the Claude Code interface, June–July 2026 — operating under the author’s direction. The tool was used to: draft and revise prose from the author’s source materials and the published record; structure and reorganise sections; convert between document formats; and cross-check quoted statistics against the published dataset and grading ledger. The reason for use is stated plainly: a single-author work of this scope benefited from AI-assisted drafting and consistency checking, and the record’s verification layer (§2.2; Data availability) is designed precisely so that no claim depends on trusting the drafting process, human or machine. All quantitative results derive from the published, independently recomputable data pipeline (dataset DOI 10.5281/zenodo.20630257), not from the AI tool; every factual claim was reviewed and verified by the author, who takes full responsibility for the content of this manuscript.
Source code available from: https://github.com/vijayjyotish/verify-jyotint
Archived software available from: https://doi.org/10.5281/zenodo.21314097
License: MIT (OSI-approved) for the software (the zero-dependency verifier verify-jyotint.mjs, the statistics reproducer reproduce-paper.mjs, and the CI workflow); the bundled sealed-forecast corpus and record artifacts are licensed CC-BY-4.0 (see LICENSE-DATA in the repository).
The verifier recomputes every SHA-256 seal, the OpenTimestamps Bitcoin anchor, the frozen grading ledger, and every statistic reported in this article from the frozen public inputs, with no dependencies beyond Node.js > = 18; a continuous-integration workflow re-verifies the record on every push.
The complete sealed corpus and graded record are openly available:
• Zenodo (dataset of record): JYOTINT Sealed Forecast Corpus, DOI 10.5281/zenodo.20630257 (concept DOI resolving to the current version), licence CC-BY 4.0.16 Contains the seal manifest (seal-manifest.json), the grading ledger (grading-ledger.json), the OpenTimestamps proofs, the calibration artifact, the analyst table (CSV/JSON), and the corpus JSON Lines.
• Hugging Face mirror: vijayjyotish/jyotint-sealed-corpus (same data package, with data card and citation file).
• Live artifacts: the REST API at https://jyotishintelligence.com/api/v1/ and the flat table at /dataset/jyotint-analyst-table.csv; distribution and delivery data at /dataset/advisory-impressions.json and/api/v1/delivery-log.json (each delivery row a public post URL, re-checkable from platform-native metadata); the regrade kit at/regrade publishes every input needed to re-score the record under a reader’s own verdicts.
• Air-gapped verification capsule: a self-contained archive (manifest + OpenTimestamps proof + zero-dependency verifier) at https://jyotishintelligence.com/dataset/crocus-capsule.zip, suitable for fully offline review on a machine with no network connection.
The record’s methodology and doctrine are documented at monograph length in The Forecasting Protocol, a three-volume monograph (Vijay Jyotish LLC, 2026; ISBNs 979–8–9969481-0-9, 979–8–9969481-4-7, and 979–8–9969-7600-3), available as free full-text PDFs at jyotishintelligence.com/book and catalogued on Open Library, Google Books, Zenodo, and the Internet Archive. The sealed Russia–Ukraine subcorpus is additionally documented in a 42-language documentary film treatment (jyotishintelligence.com/film; §3.7).
None. (No colleagues, collaborators, or co-authors contributed to this work.)
Misses are drawn as open rings, at full weight. Ten launch calls carry hour-scale leads and rise barely off the line. Basemap: Natural Earth (public domain).