248 MB
2,081 files
Updated 30 days ago
Name
Size
configs
data
docs
paper
releases
results
scripts
src
tests
verification
.gitattributes1.03 kB
xet
.gitignore1.39 kB
xet
LICENSE1.83 kB
xet
README.md34.2 kB
xet
RUNLOG.md23.6 kB
xet
pyproject.toml3.06 kB
xet
README.md

Liars Are Information: six-contribution DGX research release

Current release: releases/dgx-research-20260912-v2. Browse the complete release. Prior v1 release.

A reproducible research artifact for receiver-local multi-agent aggregation under Byzantine attacks. This release retains every original contribution and the complete first DGX verification study, then adds three implemented research extensions: clean-task calibration, clone-aware evidence caps, and strictly causal online adaptation. The six main contributions below connect research questions to code, controls, simulations, analyses, evaluation and measured conclusions.

The project studies when an untrustworthy answer can still be informative. Adversarial Informativeness Pooling (AIP) estimates a channel from a receiver's local history, then chooses TRUST, DISCARD, or INVERT. Inversion places negative evidence on a peer's answer and distributes positive evidence across alternatives. This can help when wrong agents coordinate, but calibration, correlated answers, answer-space structure and adaptive attacks can break the mechanism. The repo reports those failures as part of the result.

All new DGX experiments replay existing cached model answers and use documented symbolic attackers. No new LLM inference, generated conversation or tool-using agent trial was performed. The engineering agents who implemented and independently reviewed this release are separate from the simulated swarm identities. This distinction matters when interpreting both compute cost and statistical sample size.

Start with the six-contribution crosswalk, evaluation guide, research-question results and multi-agent implementation/review record. The original README, v1 README, historical paper, original result tables and v1 measurements are retained.

Six contributions and research questions

# Contribution Research question Runnable/evidence entry point
1 — retained Receiver-local inversion, mechanism ablations and a full confusion-matrix comparator RQ1: When does inversion improve the same local trust/discard mechanism? v1 measured replay, paired evaluation, full Dawid–Skene
2 — retained, corrected Attack-coherence geometry and empirical failure boundary RQ2: What does the implemented attacker generate, and where can it defeat the gate? corrected mathematical note, selected stationary attacks, sampler tests
3 — retained, expanded Dependence, confidence provenance and reproducible low-resource evaluation RQ3: What information is independent or trustworthy, and how can the reported work be checked? 54-cache audit, evaluation guide, release review
4 — new Disjoint clean calibration and observation transfer RQ4: Can a ceiling learned from separate clean tasks reduce honest inversion while retaining useful adversary inversion? method and assumptions, measured calibration study, implementation
5 — new Clone-aware caps on repeated peer evidence RQ5: Can history-derived groups reduce duplicate amplification without discarding useful inversion? method and assumptions, measured clone study, implementation
6 — new Strictly causal online adaptation RQ6: Do cumulative or rolling histories improve adaptation when an attacker changes behavior? method and assumptions, measured temporal study, implementation

These are six research artifact contributions, with three substantive additions in v2. They are not claims of six new algorithms, six theorems or established novelty against all prior literature. The original manuscript's five bullets—mechanism, theory, design laws, protocol, and evaluation—are explicitly mapped into retained Contributions 1–3 and the new extensions in the crosswalk. The original claims ledger keeps its own C1/C2/etc. meanings; new contribution numbers do not overwrite those claim IDs.

Measured results at a glance

Completed: 432 new simulation worlds across the three extensions, retaining all 144 v1 worlds; 548 passing tests; 306,000 new task/step-method records; independently checked prediction/fit evidence. The three standard studies together took 6 minutes 14 seconds, with peak RAM below 537 MiB, two CPU cores and no GPU inference. Counts reuse tasks and include multiple methods, validation or warmup; they are not independent benchmark samples. Exact coverage, findings and artifact dictionary.

The new results are mixed: conservative calibration often suppressed useful inversion; clone caps helped several redundant-bloc conditions but harmed some honest-clone conditions; rolling history adapted well to MMLU sleepers and failed badly under a different regime switch. The repo documents those outcomes with paired intervals and saved simulations.

The retained v1 study gives a clear example of why the repo reports both gains and failures. At 50% coherent Byzantine agents and full observation, task-averaged test accuracy was:

Benchmark Receiver-only Majority Trust-only Gated AIP AIP minus receiver, paired 95% CI
MMLU 83.7% 28.7% 0.0% 89.9% +6.2 pp [+3.7, +8.8]
MATH-500 58.6% 24.2% 0.0% 48.8% −9.8 pp [−13.0, −6.4]
BoolQ 89.3% 31.2% 0.0% 0.0% −89.3 pp [−94.2, −83.9]

These intervals resample the 80 held-out tasks after averaging three fixed swarm seeds. They condition on the cached population and legacy calibration. Rounding is for display; exact values and other comparisons are in paired_comparisons.csv and the analysis report. The MMLU gain is conditional. It does not establish general robustness to an adversarial majority.

Retained v1 accuracy, with task-bootstrap intervals and matched condition panels

At 70% corruption, a validation-selected stationary gate-aware attack reduced hard-gated MMLU accuracy to 0%. Soft variants recovered only part of the loss. Each candidate attack had its own attack-specific fitted history; this is a stationary-regime test. Contribution 6 adds a separate experiment in which behavior actually changes over time against a shared past-only state.

Contribution 1: local inversion and fair mechanism evaluation

An honest receiver is an evaluation role, not information handed to the defense. It sees its own answer and the peers visible on a task. AIP estimates receiver agreement and peer coincidence from that receiver's history, then uses the same fitted state for held-out predictions. The release fixes task alignment under missing observations and clears state before refitting.

The closest mechanism ablation is gated AIP versus trust-only AIP: same receiver, history, broadcasts and requested self-vote share, with inversion enabled or disabled. Naive inversion tests the need for the gate. Receiver-only establishes whether cooperation helps a competent receiver at all. Majority, confidence weighting, the repository's SAC-style approximation and full Dawid–Skene supply distinct comparison assumptions; they should not all be described as one-factor causal ablations.

The new full-matrix Dawid–Skene comparator estimates a latent class prior and one class-conditional confusion matrix per worker. It predicts unseen task IDs using learned worker likelihoods and can use systematic anti-expertise. The retained legacy one-coin implementation falls back to majority on unseen task IDs and is not used as this comparator. Full-matrix DS is evaluated on closed label spaces, including MMLU and BoolQ; no fixed finite confusion matrix is asserted for arbitrary MATH-500 answers.

Answer to RQ1: inversion can recover useful signal relative to trust-only in the tested coherent multiclass regimes. That does not imply it beats the receiver; MATH-500 shows the distinction. Binary tasks can fail completely. The v1 contribution specification explains the comparator assumptions, and tests cover unseen-task prediction, missingness, anti-experts under specified conditions and hidden-metadata invariance.

Contribution 2: the implemented attack law and its limits

The gate-aware attacker chooses a shared lie with probability p; otherwise it draws uniformly from a pool of K wrong candidate strings. The independent branch can draw the shared lie too. Its expected Byzantine-pair collision is therefore:

q_pair(p, K) = p² + (1 − p²) / K

For closed C-class questions, K = C−1. For open answers, K is the actual finite pool size, including a one-candidate fallback. An open answer space does not imply zero chance of collision. The old theoretical axis omitted mixed-branch collisions; it is preserved in historical result files and explicitly marked as an invalid axis for that sampler.

Pair expectation, empirical Byzantine pair coincidence and the gate's receiver-conditioned selected-partner statistic are different quantities. If a receiver is wrong, agreement among its dissenters need not be wrong-answer coincidence relative to truth. The corrected agreement identity also includes error dependence:

P(answers agree) = (1−e_i)(1−e_j) + e_i e_j q + (1+q) Cov(E_i,E_j)

The maintained attenuation note states assumptions. The conditional testing note explains why equal mean coincidence or exchangeability alone does not prove indistinguishable complete histories. The deployed gate uses more than one scalar and overlapping pair counts are dependent.

Answer to RQ2: the corrected sampler law has no population collision minimum below 1/K, while held-out replay still demonstrates substantial gate-aware harm. The empirical failure is valid evidence under its stated protocol; it is not an impossibility theorem for every history-aware defense. Validation-only selection, its ascending-p tie break, exact chosen attacks and test predictions are saved in the v1 standard directory.

Contribution 3: dependence, protocol and auditable DGX work

The honest-cache audit covers 54 caches: nine models × six benchmarks, 13,500 stored predictions. It found zero duplicate task rows and zero cross-model gold-label conflicts, while retaining three missing extracted answers and six source-marked nonreproducing rows. It computes 216 aligned cross-model error correlations, explicitly leaves eight undefined values unset, and distinguishes receiver-conditioned from gold-conditioned coincidence.

The main simulations use the seven models marked arm: frozen in the roster: gemma4_31b, granite42_30b, llama32_3b, ministral3_14b, olmo3_32b_think, phi4_mini_reasoning, and qwen38_27b. These are retained artifact registry identifiers. Ten simulated identities can repeat a model's same cached answer. Repetition increases votes but supplies no fresh independent model generation.

The original confidence-surface and forged-self-report results remain under correlation results and in the upstream README. They were not rerun as new v2 inference. A log-probability value is trustworthy only if the collector/transport establishes its provenance. Merely labeling sender-supplied metadata “logprob” does not make it unforgeable. The new extensions do not use confidence as proof that a peer is honest.

Answer to RQ3: source hashes, aligned tasks, fixed identities, local fits, conditional task statistics and actual resource receipts make the tested replay auditable. They do not make copied outputs independent or verify historical bytes when the old sidecars contain no hashes. The cache audit, v1 independent review and v2 review record what was checked.

Contribution 4: clean-task calibration and transfer

Historical thresholds were estimated on a different model population, and some preserved calibration task lists overlap the available evaluation caches. The new extension creates explicit CAL / HISTORY / VALIDATION / TEST = 20 / 30 / 20 / 30% task partitions. It estimates ceilings using only local public answers on CAL tasks in an all-honest, fully observed reference swarm, then freezes them for target worlds with full or half observation. Gold labels are not calibrator inputs.

Three ceiling choices isolate the change: the frozen legacy threshold, a pooled receiver-conditioned bootstrap upper percentile, and a conservative partner-based upper ceiling. Whole tasks carry all peer events through resampling, preserving within-task dependence. The study measures both channel decisions and answer utility: fewer honest inversions can coexist with fewer useful adversary inversions. A ceiling that saturates at one disables the relevant strict-above-ceiling test.

This extends the calibration and correlation contributions with a runnable estimator, explicit task membership, uncertainty, observation-transfer conditions, six benchmarks, saved decisions and held-out paired comparisons. The clean reference assumption is part of the design. New task separation does not prove that these historical cached tasks were never inspected before.

Answer to RQ4 and detailed evidence: calibration REPORT. Read it together with the method note, calibrator, experiment script and leakage/missingness tests. The bootstrap ceilings are practical heuristics, not finite-sample familywise error guarantees.

Calibration transfer accuracy across six benchmarks

In the completed 144-world study, all 180 partner ceilings reached one. Across 48 benchmark/corruption/visibility cells, pooled calibration had 1 positive and 21 negative pointwise paired intervals versus the legacy gate; 26 included zero. These exploratory counts have no multiplicity correction. In clean fully observed MMLU, honest channel inversions fell from 10/270 to 2/270, yet MMLU accuracy averaged equally across seeds, fractions and visibility fell from 85.7% to 67.2%. Calibration conservatism and answer utility were therefore different outcomes. Full measured report.

Contribution 5: clone-aware evidence caps

A larger swarm can simply repeat the same evidence. This extension learns conservative groups of near-identical response streams from aligned, nonmissing historical observations, with a minimum support requirement. It uses complete-link grouping so a chain of pairwise similarities does not automatically join two dissimilar endpoints. Gold, model names, clone ancestry and Byzantine labels are excluded from grouping.

At aggregation time, each peer group's total weight is capped at its strongest member before the same requested self-vote parity is applied. Peers grouped with the receiver add no duplicate self evidence. The original AIP gate fit remains unchanged within each capped/uncapped comparison. Consequently, the cap has a precise aggregation-mass meaning, while clones can still change the upstream fitted statistics across different populations.

The controlled study begins with ten fixed base identities, adds 0, 5 or 15 exact replicas of an honest or Byzantine source and evaluates only original honest receivers. It saves actual population sizes and corruption fractions. Adding honest versus Byzantine replicas changes the fraction in opposite directions, so “base f” and “realized f” are both needed. Base answers, source identity and task remain paired across clone counts.

Answer to RQ5 and detailed evidence: clone REPORT, method note, cap implementation, experiment and grouping/cap tests. Exact replicas test a narrow family; strategically diverging Sybils and independent generations remain outside the claim. Similar honest answers do not prove a common malicious owner.

Clone cap accuracy across source type and benchmark

The added-copy control shows a conditional benefit: on BoolQ at base f=0.3/full visibility, adding 15 Byzantine-source copies changed native accuracy from 86.9% to 0%, while the capped method stayed near 89.8% to 89.7%. The paired difference-in-differences was +86.8 pp [79.5, 93.2]. Conversely, 15 honest copies at base f=0.5/full visibility gave 70.4% native versus 63.9% capped accuracy on MATH-500, a −6.5 pp [−8.5, −4.6] cap effect. At other fractions, a cap advantage already exists before extra copies are added. Zero-copy controls and additional-copy analysis.

Contribution 6: adaptation without future observations

The historical window implementation is retained for provenance. The new online extension uses a separate two-phase receiver API: predict the current task, then commit its public observation. A fit can use only previously committed task positions. Tests check prefix invariance: changing future broadcasts or evaluator labels cannot change past predictions.

Four AIP retention policies share the same channel rule: the fixed initial history, the cumulative past, a 32-task rolling history, and a 64-task rolling history. Cumulative and rolling methods refresh every eight observations. Self and majority are comparison methods. The study uses 40 warmup tasks followed by 160 unique evaluation tasks and predeclared sleeper, regime-switch and toggle schedules.

The same task chronology also runs fixed honest/coherent/independent regimes. These controls help distinguish attack changes from changes in task difficulty. Chronological predictions, exact fit-task membership, fitted channel states and public traces allow a reviewer to check which evidence existed at each decision. Within-regime paired block intervals preserve local temporal dependence; the finite horizon does not support a general online guarantee.

Answer to RQ6 and detailed evidence: online REPORT, method note, causal API, experiment and causality tests. Refreshing an uninformative statistic need not restore accuracy, and a short window can lose the support needed to detect coordinated channels.

Causal online performance under the sleeper schedule

For MMLU sleepers at f=0.5, rolling32 reached 83.8%, versus 35.5% for fixed40 and 82.0% for self. Rolling32 minus fixed40 was +48.2 pp [31.5, 62.0], while its interval versus self included zero. In the coherent-to-independent schedule at f=0.7, rolling32 instead lost 57.8 pp [41.1, 68.4] relative to fixed40. These paired 200-draw intervals use seed and regime-local 16-task block resampling. Adaptation is useful in some schedules and harmful in others. Detailed temporal results.

Evaluation protocol and saved simulations

Layer Training/observation protocol Evaluation scope Evidence
Retained v1 replay 80 HISTORY / 40 VALIDATION / 80 TEST tasks; static receiver-local fits MMLU, BoolQ, MATH-500; 144 worlds; mechanism, visibility, homogeneous controls and stationary adaptive attacks standard run, additional analyses
New calibration Four disjoint partitions; clean CAL thresholds frozen before target HISTORY/TEST Six benchmarks, clean-reference transfer and target corruption/visibility conditions calibration results
New clones Historical stream grouping and matched baseline fit; frozen base broadcasts/source Three benchmarks; added exact replicas of honest or Byzantine sources clone results
New online Predict-before-update; shared chronological tasks; only committed past in a fit Three benchmarks; changed and fixed attack schedules online results

Each ordinary task accuracy is averaged over the eligible honest receivers before comparisons. Fixed seeds are also averaged within a task for paired static inference. Peer pairs, receivers, repeats, clone identities and experiment rows are not independent benchmark examples. The online study uses its documented temporal block unit. Different studies have different task splits and fitted states; their accuracies should not be pooled into one leaderboard.

The symbolic attack generator can know gold, while the defender cannot. The scorer and hidden-label diagnostics are separate from the public defense input. MATH-500 uses the repository's approximate math-equivalence scorer; the calibration experiment also identifies its GSM8K numeric scorer. Honest answer strings are preserved. Gold-assisted semantic equivalence is not provided as a defense oracle. Exact-string sensitivity is recorded where relevant.

Simulation files contain decisions and observations, not invented conversations. v1 saves 432 representative task traces and 19,935 receiver predictions. The new studies add calibration/channel traces, grouping and weight diagnostics, and complete chronological online traces. Exact artifact names and row counts are indexed in the results overview and each study's manifest.

import gzip
import json
import pandas as pd

# Inspect a retained task-level experiment and one saved public decision trace.
tasks = pd.read_parquet("results/dgx/standard/per_task.parquet")
print(tasks.columns.tolist())
print(tasks.head())
with gzip.open("results/dgx/standard/traces.jsonl.gz", "rt") as stream:
    print(json.loads(next(stream)))

# The streaming release saves every prediction and its chronological context.
online = pd.read_csv("results/extensions/online/per_step.csv")
print(online.head())

Run on a DGX with low resources

The measured host is an aarch64 DGX with NVIDIA GB10, 20 logical CPUs and roughly 128 GB system memory. The new replay deliberately uses a small CPU budget and no GPU inference. The legacy H100 campaign and its model-generation costs remain historical artifacts.

Use Python 3.12 and the recorded CPU environment. Do not install model weights or the GPU extra to run these experiments.

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements-dgx.lock.txt
PYTHONPATH=src .venv/bin/python -m pytest -q

# Retained baseline replay: choose fresh directories.
.venv/bin/python scripts/dgx_run_bounded.py --profile smoke --out results/dgx/my_smoke --timeout 120
.venv/bin/python scripts/dgx_run_bounded.py --profile standard --out results/dgx/my_standard --timeout 600

# Three new research studies. Each script handles its own resource controls.
.venv/bin/python scripts/extension_calibration.py --profile standard --out results/extensions/my_calibration
.venv/bin/python scripts/extension_clone.py --profile standard --out results/extensions/my_clone
.venv/bin/python scripts/extension_online.py --profile standard --out results/extensions/my_online --resamples 200

The extension scripts share a lock so standard studies run serially. They use two-core affinity, one native computational thread, nice 10 and hidden CUDA devices, with a 600-second per-study wall budget and documented memory controls. A sampled RSS guard differs from a hard address-space reservation; each receipt states the actual mechanism. Runtime is measured after lock acquisition. Use a new output directory and require both a completed manifest and successful resource receipt. Never overwrite supplied canonical results for a trial run.

New study Worlds Wall time Peak RAM Memory control
C4 clean calibration 144 66.12 s 486.93 MiB 2 GiB sampled RSS guard
C5 clone evidence caps 180 99.96 s 536.28 MiB 2 GiB sampled RSS guard
C6 causal online 108 207.94 s 423.42 MiB 2 GiB address space

The sum covers standard experiment execution only. Smoke repetitions, tests, independent audits, implementation and publication are separate. Machine-readable study/resource registry.

The v1 standard study took 75.76 seconds, with 504.20 MiB peak sampled RSS. Its separate 54-cache audit took 1.87 seconds and 252.8 MiB peak RSS. New extension costs appear above and in their receipts. These are costs of processing cached answers, not costs of generating model outputs.

Repository map

Location Purpose
src/aip/ Original package plus corrected static AIP and full-matrix comparator
src/aip/extensions/ Three new independently testable research implementations
scripts/ Historical phases, v1 DGX replay/audit, and bounded v2 study runners
tests/ Original suite, v1 regressions and new calibration/grouping/causality checks
data/cache/ Preserved greedy honest model outputs
data/cache_t07/ Preserved temperature/resample outputs
data/cache_adversarial/ Preserved historical model-driven attack caches
results/dgx/ Complete retained v1 main/smoke/cache audit, extra analyses and review evidence
results/extensions/ Three new canonical studies, figures, traces, evaluation logs and combined index
docs/CONTRIBUTIONS_V2.md Original-five-to-current-six crosswalk, questions and claim boundaries
docs/RESEARCH_PLAN_V2.md Internal design record, documented assumptions and evaluation rules
docs/EVALUATION_GUIDE.md Data contract, baselines, metrics, statistical units and interpretation
docs/MULTI_AGENT_V2.md Engineering-agent ownership, review findings and integration disposition
paper/ Preserved historical manuscript, references, figures and generated numeric macros
provenance/v1/ Prior release README, identity and manifest retained byte-for-byte
RELEASE.json, FILE_MANIFEST.json Current release identity and complete content fingerprints

Download and verify the Hugging Face release

The supplied location is a Hugging Face storage bucket, not a Git repository or dataset-viewer repository. Download it with the bucket API/CLI. The retained YAML front matter supports future dataset-card reuse; its historical train file labels do not define the research splits. Research split membership is saved in experiment-specific artifacts. See the official bucket documentation.

hf buckets sync hf://buckets/Nabidnur/Liars_Are_Information-bucket/releases/dgx-research-20260912-v2 ./liars-information
cd liars-information
python3 scripts/verify_research_release.py

The dated v2 release is additive. It retains the previous dated release and original source objects. The bucket root README is an index to versioned content. A file manifest fingerprints the complete new release, excluding its own bytes to avoid a recursive hash. The verification script checks that manifest, required evidence and unchanged prior scientific artifacts. The upload receipt records remote file/size checks and representative SHA-256 readbacks. No authentication credential belongs in the release.

Limits, provenance and licensing

The original manuscript and old figures remain historical evidence. Some theoretical claims and the gate-aware q axis require correction; their presence does not make them validated by this release. The maintained notes, README and measured reports state the new claims. Original paper integrity tests establish structure and macro consistency, not scientific correctness. This artifact is ready to inspect and reproduce; a paper submission must reconcile historical prose with the corrections and new evidence.

Historical calibration overlap still limits v1 comparisons. New clean-task calibration relies on a clean reference and does not erase prior exposure to this cache. Binary coordination is degenerate, open-answer scoring is approximate, and grouping exact replicas covers a narrow dependence pattern. Short causal histories may make a gate less reliable. Empirical task intervals are conditional and exploratory. Live model-driven, interactive and unrestricted adaptive attacks remain outside the new evaluation.

The original authors' LICENSE is preserved. Code and derived artifacts have the stated license; benchmark corpora and model outputs remain subject to their upstream terms. This release claims no authorship of the original algorithm, caches or manuscript. Full-matrix error modeling builds on Dawid and Skene (1979); the implemented comparator and its assumptions are documented explicitly. The contribution note records further relevant foundations without claiming their guarantees for this experiment.

Verified publication receipt. The original non-index bucket objects remain unchanged.

Total size
248 MB
Files
2,081
Last updated
Sep 11
Pre-warmed CDN
US EU US EU

Contributors