Zeno Fragility Beta
An experimental fragility score: which confident forecasts look most likely to break.
Most forecasting tools add one more forecast. Zeno does something different: it is a model of forecasters. It watches existing forecasters β weather models, epidemic hubs, economists, markets, AI forecasters β and learns when a confident forecast is about to break, by comparing it with how similar forecasts failed before. Surprise is always measured against a named reference: the crowd, the forecaster's own past, or its peers.
This is our first public release: a compact, fast, explainable model that runs on any laptop (no GPU), built to grow as more forecast history arrives.
Highlights
- Prediction markets: the 10% of confident forecasts Zeno flags as most fragile fail 1.3Γ more often than the average confident forecast (warning score 0.67 vs 0.37 for distance-from-crowd).
- Crypto: the 10% of confident forecasts Zeno flags as most fragile fail 1.8Γ more often than the average confident forecast (warning score 0.57 vs 0.44 for distance-from-crowd).
- RSV: the 10% of confident forecasts Zeno flags as most fragile fail 2.3Γ more often than the average confident forecast (warning score 0.64 vs 0.54 for distance-from-crowd).
- Planted-signal tests pass: Zeno finds a hidden skilled forecaster (+7.5%), corrects an overconfident crowd (+4.6%), and correctly does nothing when there is no signal.
- Blind spots made visible: in physical sensing, 21β24 weather models behave like only about 4 independent forecasters β exactly the shared blind spots where confident failures come from.
- Built on 2.0M forecast records across 9 sources, with data frozen and fingerprinted before training, time-ordered splits and leak checks.
Fragility warning by source
Score from 0 to 1 = how well Zeno ranks confident forecasts that later fail (0.5 is a coin flip). "β₯" is the low end of the 90% likely range. "Top-10% lift" = how much more often the most-flagged forecasts fail than the average confident forecast.
| Source | Zeno warning | Crowd-distance rule | Top-10% lift | Status |
|---|---|---|---|---|
| Prediction markets | 0.67 (β₯ 0.60) | 0.37 | 1.3Γ | ranks better than both simple rules |
| Crypto | 0.57 (β₯ 0.55) | 0.44 | 1.8Γ | beats the crowd-distance rule |
| RSV | 0.64 (β₯ 0.59) | 0.54 | 2.3Γ | beats the crowd-distance rule |
| ECB economist surveys | 0.41 (β₯ 0.38) | 0.18 | 0.3Γ | not better than chance yet |
| ForecastBench questions | 0.82 (β₯ 0.56) | 0.70 | 5.2Γ | candidate |
Weather, flu and COVID join in the next release, now that their peer-relative labels are available.
Crowd corrector (served only where it beats the plain average on unseen confirmation data)
| Source | Gain over the plain crowd average |
|---|---|
| Physical sensors (air, urban, sea, agriculture) | +18.6% (90% range +15.5% to +22.2%) |
A safety rule keeps the plain crowd average for every other source until the corrector beats it on both test and confirmation data.
Model details
| Type | Per-source gradient-boosted trees + logistic fallback (scikit-learn) |
| Parameters | 62,880 learned numbers (tree split thresholds + leaf values + logistic weights, 6 sources) β stored in model.safetensors |
| Formats | ONNX (onnx/<source>.onnx, verified identical to the original to 1e-7), safetensors, scikit-learn pickle |
| Inputs (6) | crowd distance, long and recent track record vs crowd, crowd spread, crowd size, own confidence |
| Output | experimental fragility score, 0β1 (higher = more fragile; a ranking, not a calibrated probability) |
| Fitting rows | 38,459 confident forecasts (served models) |
| Training data | Offdiagonal forecast corpus, 2.0M records, 9 sources, frozen snapshot ep-20261001T0533Z-v035 |
| Run | v038-full-20261001T1537Z (CPU only) |
Download and use
pip install huggingface_hub scikit-learn numpy
hf download ZenoDivergent/zeno-fragility-beta --local-dir zeno-fragility-beta # or: python -c "from huggingface_hub import snapshot_download as d; d('ZenoDivergent/zeno-fragility-beta', local_dir='zeno-fragility-beta')"
import sys; sys.path.insert(0, "zeno-fragility-beta")
from zeno_fragility import load, crowd_features, warn
m = load("market") # available: crypto, ecb, forecastbench, health_rsv, market, physical
f = crowd_features(own=0.82, crowd=[0.55, 0.60, 0.48, 0.82, 0.58], own_confidence=0.64,
record_long=(12.1, 14.0), record_recent=(3.0, 3.1), m0=m["prior_track_record_m0"])
print(f"fragility score: {warn(m, f):.2f}")
Run with ONNX (any language, no scikit-learn needed):
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
sess = ort.InferenceSession(hf_hub_download("ZenoDivergent/zeno-fragility-beta", "onnx/market.logistic.onnx")) # each source's selected model: see reports/per_source.json β served
# features in this order: crowd_distance, log_record_long, log_record_recent, crowd_spread, log_crowd_size, own_confidence
x = np.array([[2.1, -0.15, -0.03, 0.9, 1.8, 0.64]], dtype=np.float32)
print(sess.run(None, {"features": x})[1][0, 1]) # fragility score
Why is it so small?
The served models were fitted on 38,459 confident forecasts with known outcomes across 6 sources, reading 6 signals per forecast. Small means fast, explainable and hard to overfit. The next model reads whole forecasting events β every forecaster's range, track record and shared blind spots β and this beta is the baseline it has to beat.
Files: onnx/<source>.onnx (onnx/<source>.logistic.onnx where logistic is selected), model.safetensors, config.json, warning/<source>.pkl (fragility model per source), correctors/<source>.pkl (served correctors), zeno_fragility.py (helper), reports/per_source.json (every number, including the ones that didn't pass).
Limits
- Research beta, early stage; scores are per source and are rankings, not calibrated probabilities.
- Each source uses the model selected on held-out data (stored as
selectedin each file;warn()uses it by default).
Links
zenodivergent.dev Β· run reports in dataset mbarbosa1/zeno-divergent-runs under v038/v038-full-20261001T1537Z/
- Downloads last month
- 68