NoFutureData

Evaluation and falsification

Date: 2026-09-17 (Asia/Taipei) Evaluation update: GPT-5.6 Sol (gpt-5.6-sol)

NoFutureData asks a narrow research question: can a small repository-local guard make temporal causality assumptions testable without requiring a feature store, backtesting engine, or ML framework?

The project does not treat a single benchmark score as proof. It separates three forms of evidence because they fail in different ways.

Evidence layer What it tests Main failure mode
Static source rules Known future-looking source patterns Misses arbitrary indirect dependencies and can be context-sensitive
Behavioral invariance Whether past outputs change after future rows are removed or mutated Requires an executable transform and representative inputs
Availability contracts Whether a record was known and eligible before a historical decision Depends on correct provider/publication semantics

Current reproducible results

1. Static conformance corpus

benchmarks/corpus.json contains 31 deterministic cases: 18 planted leaks and 13 safe controls. Every shipped semantic static rule (SRC001+), including the opt-in contextual rules, has a leaking case and a safe counter-example where applicable. SRC000 is the parse/configuration failure sentinel and is tested separately rather than treated as a leakage pattern.

Current result:

These are conformance metrics on a corpus designed around the shipped rules. They are not estimates of real-world scanner recall.

2. Jupyter notebook fixture corpus

notebooks/fixtures/ contains 24 minimal notebooks: exactly one planted leak and one paired safe control for each of the 12 shipped semantic SRC001+ rules. The fixtures are generated deterministically from notebooks/fixtures/manifest.json and validated by benchmarks/run_notebook_corpus.py.

Current notebook result: 24/24 cases pass across 12/12 paired semantic rules. For every leaking notebook, CI checks the exact rule ID and notebook cell number; for every safe notebook it requires zero findings. The runner also verifies that the checked-in notebook JSON still matches the manifest source, that finding metadata points back to the correct notebook path, and that the contextual SRC011/SRC012 leak examples remain clean without declared time-series context.

The notebook corpus is a deterministic interface/conformance test. It does not measure how often temporal leakage occurs in arbitrary notebooks or estimate real-world scanner recall.

3. Static vs behavioral method ablation

benchmarks/method_comparison.py contains six cases: four leaks and two safe controls. Two leaks use explicit source patterns known by the AST scanner; two use indirect future dependencies that intentionally sit outside its rule set.

Method Precision Recall Specificity F1
Static rules 1.000 0.500 1.000 0.667
Prefix invariance 1.000 1.000 1.000 1.000
Future-mutation invariance 1.000 1.000 1.000 1.000
Runtime union 1.000 1.000 1.000 1.000
Static + runtime 1.000 1.000 1.000 1.000

This experiment is designed to expose method boundaries. It demonstrates that the behavioral tests catch indirect future dependence that the current static rules do not. It does not imply that behavioral testing has perfect recall on arbitrary real pipelines.

4. Synthetic downstream metric inflation

benchmarks/downstream_metric_inflation.py connects the source-level problem to a model-level consequence. A fixed-seed AR(1) process produces 2,400 forecasting rows. Both models use the same ordinary least-squares implementation and the same chronological 1,600/800 train/test split. The causal model sees lag/current observations; the leaking model receives one additional feature: the next latent observation, which is unavailable at the forecast time.

Model Holdout R² Holdout RMSE
Causal lag/current features 0.630 0.962
Adds unavailable future feature 0.991 0.152

The R² inflation is +0.361 and the leaked RMSE is 0.158× the causal RMSE. All 7/7 acceptance checks pass: the split is non-empty, the causal model stays non-trivial, the leaked model becomes near-perfect, the performance gap is material, the safe lag source remains clean, and the future-shift source emits SRC001.

This is a synthetic mechanism demonstration, not an estimate of how much leakage improves metrics in real projects. Its purpose is to make the practical failure mode reviewable while keeping the causal claim deliberately narrow.

5. Mutation intervention validity

benchmarks/mutation_validity.py tests four materially different input contracts: bounded probabilities, valid OHLC candles, simplex-normalized vectors, and monotonic cumulative counters. For probabilities, the generic numeric mutator intentionally leaves the valid domain, so the experiment must reject that intervention rather than interpret the resulting behavior as leakage. For OHLC, a deliberately invalid mutator makes close > high and must also be rejected. For normalized vectors, a deliberately invalid mutation breaks the non-negative unit-sum contract and must be rejected. For cumulative counters, a future value that drops below its historical predecessor must be rejected. Domain-preserving mutators are then used on paired causal and leaking transforms in all four domains.

Current result: 13/13 checks pass.

These four input domains are not proof that the default intervention is valid for arbitrary schemas. They establish an executable contract for testing domain-aware counterfactuals without interpreting malformed interventions as leakage evidence.

6. Intervention sensitivity

benchmarks/intervention_sensitivity.py asks whether the causal/leaking conclusions survive more than one hand-picked counterfactual. It evaluates all four intervention domains at three historical cut points and two valid mutation strengths. Each scenario contributes two checks: the causal control must remain clean and the planted future dependency must still be detected.

Current result: 48/48 checks pass.

The grid is intentionally fixed and reviewable. A failure at any valid cut point or mutation strength is evidence against the current behavioral claim; the benchmark does not select or report only the best-performing intervention.

7. Real-pandas behavioral transfer

benchmarks/behavioral_generalization.py removes one simplifying assumption from the method ablation: its transforms call pandas directly instead of using hand-written equivalents. It now covers ordinary trailing transforms, interleaved grouped/stateful transforms, and irregular timestamp/missing-data pipelines plus index alignment, resampling boundaries, and multi-column stateful pipelines with paired causal and future-dependent cases.

Behavioral-transfer result: 18/18 executable pandas cases match the declared causal labels.

This transfer result reduces dependence on benchmark-authored surrogate implementations, but the 18 operations are still curated and finite. It does not establish recall over arbitrary pandas pipelines or other dataframe systems.

8. Revision/vintage generative robustness

benchmarks/revision_vintage_robustness.py uses fixed seed 20260916 to generate 24 different append-only vintage histories with explicit positive eligibility delays. Each generated history is checked five ways: unique vintages in their original order must pass, the same rows in a shuffled order must still pass, one injected duplicate known_at for the same logical observation must emit REV001, a decision between known_at and eligible_from must not expose that vintage early, and a decision just after eligibility must select the latest eligible vintage identically even when right-side rows are shuffled.

Current result: 120/120 checks pass across 24 generated trials.

The fixed seed makes failures exactly reproducible while varying series count, observation count, vintage count, values, eligibility delays, timestamps, and row order. This is a deterministic generative sample, not a formal proof over all revision histories.

benchmarks/property_revision_vintage.py lifts the five fixed-seed invariants plus seven multi-column/null edge invariants into Hypothesis strategies. Unlike the fixed-seed sample, Hypothesis varies the history shape, row permutation, multi-column key partitions, and null/revision interactions while retaining shrinking, so a failure is reduced toward a smaller reproducible counterexample.

Property-based result: 768/768 invariant evaluations pass across 64 Hypothesis examples.

The property strategy is deliberately bounded. It strengthens counterexample search beyond the 24 fixed-seed histories, but it remains finite evidence rather than a proof over all revision processes.

10. Documentation-backed external reproductions

benchmarks/external_reproductions.json preserves examples derived from public Freqtrade, pandas, scikit-learn, Polars, Xarray, NumPy, Dask, PySpark pandas, and Snowpark pandas documentation, including source URLs, context requirements, and the reviewed current detector behavior.

Current result:

Three earlier misses are now covered with deliberately narrow rules: SRC008 flags a direct whole-series aggregate only when it is assigned back to a column on the same dataframe, SRC009 flags negative absolute iloc positions, and SRC010 flags fixed/day resample aggregations that use the left interval label. Nine source-only misses are intentionally kept visible and are recovered only after the caller declares time-series semantics. scikit-learn documents that classical folds are inappropriate for time-series evaluation, but the same KFold, cross_val_score(..., cv=5), or nested integer-CV search/evaluation source can be legitimate for IID data. With the explicit temporal_context="time_series" contract, SRC011 detects both single-level cases and both levels of the nested GridSearchCV + cross_val_score reproduction. It also covers GroupKFold and GroupShuffleSplit in declared panel/time-series evaluation: official scikit-learn documentation states that GroupKFold groups appear in arbitrary fold order and that GroupShuffleSplit uses randomized group partitions, neither of which alone establishes chronological train-before-test ordering. The metadata-routing-capable learning_curve, validation_curve, and permutation_test_score APIs are also context-gated when cv is omitted, None, or a literal integer because their documented default splitters are KFold/StratifiedKFold. Paired explicit TimeSeriesSplit controls remain clean. NumPy’s official roll semantics provide a separate contextual case: np.roll(values, -1) moves later elements toward earlier positions and therefore creates a future dependency on ordered data. SRC012 detects only the narrow np/numpy attribute-call form with a literal negative shift under declared time-series context. Source-only scans, arbitrary .roll methods, dynamic shifts, and the paired positive-roll lag control remain clean.

PySpark pandas and Snowpark pandas provide two additional backend-transfer checks. Their official Series.shift and Series.diff documentation explicitly supports negative periods, while the paired positive-period examples use previous rows. The existing SRC001 and SRC005 rules reproduce those documented semantics without any backend-specific special case.

Xarray provides a separate transfer test for the existing shift rule. Its DataArray.shift API accepts offsets keyed by dimension, so shift(time=-1) moves later values toward earlier timestamps. SRC001 recognizes only a small set of explicitly temporal dimension names; shift(time=1) and a non-temporal shift(axis=-1) control stay clean rather than generalizing every negative keyword argument.

11. Public GitHub-reported cases

benchmarks/reported_cases.json adds a separate provenance class: public issue reports from real projects rather than API documentation or examples authored for NoFutureData. The runner preserves the reviewed detector behavior even when that behavior is a known limitation.

Current reported-case result: 3/3 reviewed baselines match.

The leak reports are preserved as reporter claims. A GitHub issue is evidence of real user pain and a reproducible stated contract; it is not automatically proof that the upstream maintainer accepted the diagnosis. The corpus therefore tests NoFutureData’s response to the reported contract and keeps a real false-positive boundary visible instead of converting issue counts into recall/precision claims.

Falsification contract

The project treats the following outcomes as evidence against its current design, not as cases to hide:

Release reproducibility is checked from the built source distribution as well as the repository checkout: CI requires the benchmark runners, research brief, evaluation note, schema, examples, and fixtures to survive packaging, then runs this evaluation suite from the extracted sdist with PYTHONPATH=src.

Any such case should be reduced to a minimal reproduction and committed before changing the detector. The reproduction then becomes a regression case.

Reproduce locally

python -m pip install -e ".[pandas,dev]"
python benchmarks/run_evaluation.py
python benchmarks/run_benchmark.py
python benchmarks/method_comparison.py
python benchmarks/mutation_validity.py
python benchmarks/intervention_sensitivity.py
python benchmarks/behavioral_generalization.py
python benchmarks/revision_vintage_robustness.py
python benchmarks/property_revision_vintage.py
python benchmarks/run_external_reproductions.py
python benchmarks/run_reported_cases.py
pytest
nofuture audit-manifest examples/temporal_contract.json
nofuture scan src tests examples/safe_pipeline.py benchmarks
nofuture scan path/to/time_series_project --time-series

The pandas extra is required by the generative revision/vintage experiments so they can exercise the shipped point_in_time_join implementation. Hypothesis is developer-only and is used for counterexample generation/shrinking; the core package still has no mandatory third-party runtime dependency.

run_evaluation.py --json emits one report containing all eight evaluation components, the package version, UTC generation time, current miss count, context-required miss count, source-project list, and explicit claim limits. Its exit status is based on reviewed benchmark contracts rather than requiring the intentionally incomplete external detector to reach 100% recall.

Developer dependency audit

Audit date: 2026-09-17 (Asia/Taipei).

External provenance

The external corpus is curated and non-random. Its rates describe only the checked reproductions and must not be presented as population-level recall, precision, or prevalence estimates.