Mehmet Demir Güven

Mehmet Demir Güven

Computer Science, ETH Zürich

I build machine learning and quantitative finance systems end to end: data pipelines, libraries and experiments, and the proofs that decide whether to trust them.

Education

ETH Zürich. BSc Computer Science2026 to 2029

First-year coursework: Linear Algebra, Discrete Mathematics, Data Structures and Algorithms, Introduction to Programming, Analysis I, Algorithms and Probability, Parallel Programming, Digital Design and Computer Architecture.

Turkish (native), German (C2), English (C1)

Selected work

Each project opens into a full write-up: the mathematics, live simulations, and the engineering behind it.

Machine learning

Symmetry in sine networks

I built a corpus of about 1.8 million fitted sine networks and an evaluation ladder for models that read their weights, then proved the complete function-preserving symmetry group, D∞ ≀ Sn per layer. That group alone costs 79.1 of 80.4 accuracy points; the exactly invariant reader I designed recovers 0.917 of them and ships as the phasorkit library.∎

PyTorch library289 tests, including property tests on the group actionpreregistered
Run-to-run variance in RL post-training

I designed a framework-agnostic forecasting API, with an exact backend and an MLX backend for pretrained models, that predicts the across-seed spread of RL post-training from one stored run through a covariance recursion and an adjoint pass: 1.11x median error over 192 prospective settings, with its failures on Qwen2.5 measured and published.

PyTorch and MLX backendsexactness testsprospective evaluation

Quantitative finance

Cross-asset price impact is set-identified

I proved that the confounding gap between estimated and structural cross-impact has rank at most K + rank(B) and that its cost error is a low-rank quadratic form, then built a sharded known-truth simulator, 107 observations in 100 frozen shards with byte-for-byte replay, that verifies every derivation: an index basket is mispriced by 54.23%, a dollar-neutral one by nothing.∎

strict mypy408 testssharded simulation, byte-for-byte replay
The hidden options market

I engineered a validated ingestion pipeline for the OCC's daily clearing reports, producing a 13-month panel of the unlisted FLEX market (9.06m series-days, $596bn of mark value), with leg grouping that corrects 49x notional errors. A placebo-tested fixed-effects regression shows the hidden inventory does not predict next-day volatility.

validated ingestion pipeline256 testsfree data only
Nothing beats zero

I built a known-truth Monte Carlo harness, in Python with a Rust indicator library, that measures the real size of forecast-comparison tests: at a nominal 5%, Diebold-Mariano on panel rows rejects a true null 79.13% of the time on 100 correlated stocks. I derived the closed form, rebuilt the apparatus, and made the whole run replay byte for byte.∎

Python with a Rust core606 testshash-verified replay
Crypto returns and the martingale null

I built a leak-proof walk-forward evaluation, with purging, embargo and in-fold scaling enforced by tests, and used it to test 18 machine-learning settings for BTC, ETH and SOL with the correct test for nested models. 5 reject no-predictability, 2 survive false-discovery control, none survive family-wise control or costs.

154 tests, 96% coverageCI on Python 3.12 and 3.13

Notebook

Probability, in the margins

Eight theorems about random walks and Brownian motion, each simulated in the browser until it happens: Donsker's invariance principle, the arcsine law, the law of the iterated logarithm, optional stopping, quadratic variation, Itô against Stratonovich, Girsanov's change of measure, and Black-Scholes as a Feynman-Kac expectation. The same processes run quietly through this page.

Toolbox

Languages
Python, C++, Rust, TypeScript, Java, Swift, SQL
ML and numerics
PyTorch, MLX, NumPy, SciPy, Polars, pandas, scikit-learn, Hugging Face Transformers
Engineering practice
Property and exactness tests, strict typing, continuous integration, preregistered experiments, byte-for-byte reproducible runs

Olympiads

TÜBİTAK National Informatics Olympiad, 2025. Qualified as one of 511 from 18,463 students.

TÜBİTAK National Mathematics Olympiad, 2024. Advanced to the proof-based national round.

Symmetry in sine networks

Parameter symmetry and weight-space learning on sine-network implicit neural representations

July 2026 to present. Manuscript in arXiv moderation.

StackPyTorch
Tests289
Corpus1.8M networks
Commits251

Start here. A real sine network, one input, two hidden layers of four neurons, one output, evaluated in your browser. Every button changes the parameters; none changes the function on the right. This page proves those are the only such changes, and measures what they cost a model that reads weights.

Abstract. An implicit neural representation stores a signal as the weights of a small network fitted to it. Classifiers that read those weights work when every network starts from a shared initialization and collapse when each starts from its own: on MNIST, 80.4 accuracy points separate the two, with the same images, architecture and reader. I characterise the complete function-preserving symmetry group of sine networks, show by intervention that it alone reproduces 79.1 of the 80.4 points, and build a reader that is exactly invariant to it.

1. Setup

Each image is stored as a SIREN with two hidden layers of width 32, fitted to map a pixel coordinate to its intensity:

\[ \Phi_\theta(x) = W_3 \sin\!\big(W_2 \sin(W_1 x + b_1) + b_2\big) + b_3, \qquad \theta = (W_\ell, b_\ell)_{\ell=1}^{3}. \tag{1} \]

I fitted about 1.8 million such networks across MNIST, FashionMNIST and CIFAR-10, on one laptop, under four protocols that differ in exactly one source of variation. A corpus is admitted only if a classifier trained on renders of its networks matches one trained on real pixels, so no result can come from information lost in fitting.

This is not a picture. It is a SIREN with two hidden layers of 80 neurons, fitted to the text of my name and evaluated pixel by pixel in your browser, assembled here one neuron at a time. Now apply a random element of the group \(G\) from Section 2 to all of its parameters.

An INR you can inspect. Thousands of parameters change and not one pixel does, up to floating-point rounding: Theorem 1 says that for a generic network this is the only way it can happen.

2. The symmetry group

Two identities of the sine do all the work:

\[ \sin(-z) = -\sin z, \qquad \sin(z + \pi) = -\sin z. \tag{2} \]

For neuron \(i\) of a hidden layer, with incoming row \(w_i\), bias \(b_i\) and outgoing column \(v_i\), they give two function-preserving moves:

\[ \sigma_i : (w_i, b_i, v_i) \mapsto (-w_i, -b_i, -v_i), \qquad \tau_i : (b_i, v_i) \mapsto (b_i + \pi, -v_i). \tag{3} \]

Since \(\sigma_i^2 = 1\) and \(\sigma_i \tau_i \sigma_i = \tau_i^{-1}\), each neuron carries a copy of the infinite dihedral group, and reordering the neurons of a layer adds a wreath product:

\[ \langle \sigma_i, \tau_i \rangle \cong D_\infty, \qquad G = \prod_{\ell} D_\infty \wr S_{n_\ell} = \prod_{\ell} D_\infty^{\,n_\ell} \rtimes S_{n_\ell}. \tag{4} \]

Theorem 1 (completeness)

For generic parameters of a sine network with one or two hidden layers,

\[ \Phi_\theta = \Phi_{\theta'} \iff \theta' \in G \cdot \theta. \]

No function-preserving change exists outside \(G\). Proof in the manuscript; the depth-two case is also checked numerically in the repository.∎

3. How much of the gap is symmetry?

Write \(a_r\) for the test accuracy of rung \(r\) of the ladder in Figure 1. The share of the gap a rung recovers is

\[ f(r) = \frac{a_r - a_{\mathrm{W3}}}{a_{\mathrm{W1}} - a_{\mathrm{W3}}}, \tag{5} \]

with \(a_{\mathrm{W1}} = 94.36\) for weights from a shared initialization and \(a_{\mathrm{W3}} = 13.92\) from independent ones.

Exact reframingStandard treatmentEnds of the gapReference
Figure 1. The decomposition ladder on MNIST. The standard treatments for symmetry barely move. Chance is 10.

The causal question is answered by intervention: applying a random element of \(G\) to every network of the shared corpus changes no function, yet costs 79.1 of the 80.4 points. A factorial design over the four kinds of change \(N = \{\text{reorder}, \text{sign}, \pi, 2\pi\}\) splits that loss by Shapley value,

\[ \phi_j = \sum_{S \subseteq N \setminus \{j\}} \frac{|S|!\,\big(|N| - |S| - 1\big)!}{|N|!}\,\big[v(S \cup \{j\}) - v(S)\big], \tag{6} \]

where \(v(S)\) is the accuracy lost when only the changes in \(S\) are applied. A scramble that does change the function costs only 0.7 to 2.0 points, so the damage comes from the group and not from reordering as such.

Reordering neuronsSign flipsπ bias shifts2π windings
Figure 2. Shapley values of the four kinds of change, in accuracy points.

4. An exactly invariant reader

Encoding each bias on the circle removes the infinite part of the group:

\[ b \mapsto z = (\cos b, \sin b) \in S^1, \qquad \tau : z \mapsto -z, \qquad \sigma : z \mapsto \bar z. \tag{7} \]

The \(2\pi\) winding acts trivially on \(z\), so on the encoded parameters the group acts through a finite group of sign changes, and a reader built on sign-invariant combinations is exactly invariant on the raw parameters. At the same parameter count it recovers 0.917 of the gap, against 0.628 for the best alignment, and scores 93.51, 74.81 and 44.20 on the published INR benchmarks: third, third and first among readers that use weights alone.

Figure 3. Share of the MNIST gap recovered, \(f(r)\) from equation (5).

Connections: diffusions, invariance and Fourier features

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

Training noise is a diffusion

\[ d\theta_t = -\nabla L(\theta_t)\,dt + \sqrt{2T}\,dW_t, \qquad \pi(\theta) \propto e^{-L(\theta)/T} \]

Near a minimum, stochastic gradient descent behaves like Langevin dynamics, whose stationary law is the Gibbs measure \(\pi\). Because the loss is invariant, \(L(g\theta) = L(\theta)\) for every \(g \in G\), and the group acts by volume-preserving maps, \(\pi\) is invariant too: \(\pi(g\theta) = \pi(\theta)\). Independent runs therefore cannot prefer one member of an orbit over another. That is the random-initialization corpus, seen as a diffusion.

Averaging over a group

\[ \bar f(\theta) = \frac{1}{|G^{\prime}|}\sum_{g \in G^{\prime}} f(g\theta) \]

The Reynolds operator turns any reader into an invariant one, but only when the group is finite. The phasor encoding is what makes the infinite group \(G\) act through a finite group \(G^{\prime}\) of sign changes, so exact invariance becomes a finite computation rather than an approximation.

Sine layers are Fourier features

\[ \mathbb{E}_{w}\big[\cos\big(w^{\top}(x - y)\big)\big] = k(x - y) \]

By Bochner's theorem every shift-invariant kernel is the expectation of a cosine, so a first layer \(\sin(w^{\top}x + b)\) with random \(w\) is a random Fourier feature map (Rahimi and Recht). A SIREN learns its Fourier basis, and the \(\pi\) shift in the symmetry group is the phase ambiguity of that basis.

5. How it is built

The pipeline. Black boxes are the gates: a corpus that fails the render check never reaches a reader, and a prediction is frozen before the experiment that tests it exists.
import torch
import phasorkit as pk

def build_corpus(n: int = 256, seed: int = 0) -> pk.INRCorpus:
    arch = pk.Architecture(dims=[2, 32, 32, 1])
    g = torch.Generator().manual_seed(seed)
    weights = torch.randn(n, arch.n_params, generator=g) * 0.3
    labels = torch.randint(0, 2, (n,), generator=g)
    weights[labels == 1] *= 2.0
    return pk.INRCorpus.from_flat(weights, arch, labels=labels)

report = pk.compare(build_corpus(), seed=0)   # probe vs weight reader, matched FLOPs

Group theory, weight-space learning, invariant architectures, Shapley decomposition, factorial experiments, preregistration, PyTorch.

Run-to-run variance in RL post-training

Forecasting the across-seed spread of reinforcement learning post-training from a single stored run

August 2026.

The idea in one picture. Forty runs of noisy gradient descent on a quadratic, differing only in their noise, simulated in your browser. The band is ±2 standard deviations predicted by recursion (2) below, computed without looking at the runs. Remove the learning signal and the contraction disappears. An illustration of the mechanism, not the paper's data.

Abstract. A score reported from RL post-training is the output of a stochastic procedure, and the only instrument the field has for its spread is to train again with new seeds. Gradient noise injected at one update is filtered by the updates that follow rather than accumulated, so the covariance of a run about its mean obeys a linear recursion. Carrying the gradient of the reported metric backwards along one stored trajectory turns that recursion into an error bar. Tested prospectively it holds to 1.11x over 192 settings; on pretrained models it fails, and I report by how much.

1. The recursion

With \(P\) prompts per update, step size \(\eta\) and a group-based advantage estimator, the parameters follow a noisy gradient step, and linearizing around the mean trajectory gives the covariance of a run about its mean:

\[ \theta_{t+1} = \theta_t + \eta\,\hat g_t, \qquad \mathbb{E}\,\hat g_t = g(\theta_t), \qquad \operatorname{Cov}\hat g_t = \Sigma_t / P, \tag{1} \] \[ S_{t+1} = A_t S_t A_t^{\top} + \frac{\eta^2}{P}\,\Sigma_t, \qquad A_t = I + \eta J_t, \qquad J_t = \nabla g(\theta_t). \tag{2} \]

When the eigenvalues of \(A_t\) sit inside the unit disc, noise injected early is filtered out by the updates that follow. Against exact Monte Carlo over twelve settings the recursion is accurate to 1.09x at the median.

Real rewardsCoin-flip rewards
Figure 1. Divergence between real runs over eighty updates, relative to its start, log scale, endpoints only. Swapping rewards for coin flips, with everything else untouched, turns contraction into growth.
Figure 2. Error of three ways to propagate the covariance against exact Monte Carlo, log scale.

2. The adjoint

Unrolling (2) and projecting onto the reported metric \(M\) gives the across-seed variance as a sum of scalars, computed by a backward pass along the one stored run:

\[ \operatorname{Var} M(\theta_T) \approx \sum_{t=0}^{T-1} \frac{\eta^2}{P}\; b_{t+1}^{\top} \Sigma_t\, b_{t+1}, \qquad b_T = \nabla M(\theta_T), \qquad b_t = A_t^{\top} b_{t+1}. \tag{3} \]

For count-based estimators the mean update is the gradient of a scalar potential, so the Jacobian is symmetric and each backward step is a Hessian-vector product; the measured asymmetry is 3e-10.

\[ g(\theta) = \nabla_\theta\, \mathbb{E}_x\!\left[\Lambda\big(p_x(\theta)\big)\right] \quad\Longrightarrow\quad J_t = J_t^{\top}. \tag{4} \]

An exact finite-\(G\) covariance for the advantage estimators in use shows they differ only by a weight on prompt difficulty, and fixes the optimal group size at \(G^{*} = 1 + \sqrt{\tau_w / \tau_b}\).

3. Evidence

One run was trained and frozen, and the forecast recorded, before the remaining seeds were trained. Over 192 settings spanning a 47x range of spreads the median absolute error is 1.11x, the worst case 1.58x, and the forecast lands inside the measured 95% interval in 162 of them. Across nine settings the log-log slope of divergence runs from −1.15 to +0.26, where a random walk needs +1.

Figure 3. Scaling exponents with 95% intervals against the predicted −1/2 in prompt count and +1/2 in drift.
Rollout samplingPrompt selection
Figure 4. Where the forecast variance comes from, by group size \(G\).

4. Where it fails

On pretrained models the forecast is short: by 25x on Qwen2.5-0.5B under plain ascent and by 3x to 17x on Qwen2.5-7B under Adam, because runs end in groups rather than scattered around one. An earlier route via \(\sqrt{\mathrm{KL}}\), with a fitted exponent of +0.10 against a predicted +0.5, is marked withdrawn rather than deleted.

Figure 5. Forecast error as a ratio of predicted to measured spread. Blue where the comparison is exact; orange on pretrained models.

Connections: Ornstein-Uhlenbeck, Itô isometry and control variates

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

An Ornstein-Uhlenbeck process

\[ dX_t = -H X_t\,dt + \sqrt{\tfrac{\eta}{P}}\;\Sigma^{1/2} dW_t, \qquad H S_\infty + S_\infty H = \tfrac{\eta}{P}\,\Sigma \]

Near an optimum with Hessian \(H\), the deviation of a run from the mean run is an Ornstein-Uhlenbeck process, and its stationary covariance solves the continuous Lyapunov equation. Recursion (2) is the exact discrete version, with \(A = I - \eta H \approx e^{-\eta H}\). The fan at the top of this page is a discretized OU process; switching off the signal sets \(H = 0\) and leaves a Brownian motion, whose variance grows linearly: the coin-flip blow-up.

Why one run is enough: the Itô isometry

\[ \operatorname{Var}\!\int_0^T b_s^{\top}\sigma_s\,dW_s = \int_0^T \big\lVert \sigma_s^{\top} b_s \big\rVert^2 ds \]

The variance of a stochastic integral is the integral of its squared integrand. Equation (3) is its discrete form: a sum over updates of the injected noise, projected on \(b_s\), how much the reported metric still depends on that update. Each term is computable along the one stored run.

Policy gradients and control variates

\[ \hat g = \frac{1}{G}\sum_{i=1}^{G}\Big(r_i - \frac{1}{G-1}\sum_{j \neq i} r_j\Big)\,\nabla_\theta \log \pi_\theta(y_i \mid x) \]

The RLOO estimator subtracts a leave-one-out baseline: a control variate that leaves the mean gradient unchanged and shrinks \(\Sigma_t\) in recursion (2). The optimal group size \(G^{*}\) trades that variance reduction against the cost of extra rollouts.

5. Build

The forecast is a protocol, not a training framework: any trainer that answers two questions about each stored update can be measured. Two backends ship, one exact for enumerable policies and one for pretrained models through MLX on Apple Silicon, using common random numbers so finite differences are directional derivatives rather than differences of noise. Every run leaves a manifest of configuration, commit and platform, and no number in the paper is transcribed by hand.

From one stored run to a scored forecast. The black box is the step no other tool provides.
from caliper.forecast import run_backward

class MyUpdate:                       # step_size, n_prompts as attributes
    def projected_variance(self, b):    # Var over prompts of b . per-prompt gradient
        ...
    def jacobian_vector(self, b):       # J b, one Hessian-vector product
        ...

result = run_backward(metric_gradient, updates)
result.std                # forecast s.d. of the reported score across seeds
result.interval(0.784)    # a normal interval around the number you report

Stochastic approximation, adjoint methods, RLVR and RLOO, Hessian-vector products, prospective evaluation, PyTorch, MLX, Qwen2.5. 133 tests and continuous integration.

Cross-asset price impact is set-identified

Spurious or structural? Low-rank confounding and partial identification of cross-asset price impact

July to August 2026. Preprint, version 0.3.

GatesG0 and G1 passed
Tests408
Simulation107 draws
Typingstrict mypy

The puzzle, simulated. Six stocks with no cross-impact at all and \(K\) common factors moving both prices and order flow. Each change draws 4,000 observations and runs \(\widehat\Lambda = \widehat\Sigma_{rq}\widehat\Sigma_{qq}^{-1}\) in your browser. The regression reports cross-impact that does not exist, and the singular values of its error show exactly \(K\) directions: Theorem 2, below, live.

Abstract. Cross-asset return-on-flow coefficients are routinely read as entries of a structural price-impact matrix. In a published study, adding one principal-component control flips the mean coefficient from +0.032 to −0.039 and the share of negative coefficients from 23.09% to 84.46%. I show that neither reading is identified, characterise exactly how, and price the consequence. Each step of the argument is one row of the derivation below.

\[ r_t = \Lambda q_t + \Gamma f_t + u_t, \qquad q_t = B r_t + \Delta_f f_t + v_t \]
Model
Returns and flows in \(\mathbb{R}^N\) move each other through \(\Lambda\) and \(B\), and both load on \(K\) latent factors \(f_t\). With \(H = (I - B\Lambda)^{-1}\), flows reduce to \(q_t = Pf_t + Uu_t + Hv_t\).
\[ \operatorname{plim}\widehat\Lambda_{\mathrm{OLS}} = \Lambda + \Gamma\Sigma_f P^{\top}\Sigma_{qq}^{-1} + \Sigma_u U^{\top}\Sigma_{qq}^{-1} \]
Theorem 1
What regression converges to: the truth plus a confounding gap \(G\) from the factor channel and the feedback channel.
\[ \operatorname{rank}(G) \le K + \operatorname{rank}(B) \]
Theorem 2
The factor channel passes through a \(K\)-dimensional bottleneck, feedback through the column space of \(B\), and rank is subadditive. Attained generically: observed ranks 3, 4, 5 meet the bound at rank(\(B\)) = 0, 1, 2.
\[ \mathcal{D}_K = \{D + R : D \text{ diagonal},\ \operatorname{rank}(R) \le K\} \]
Corollary
A diagonal truth produces an estimate confined to \(\mathcal{D}_K\), but not small: spurious off-diagonals reach 0.2207 against genuine own-impact from 0.2061 to 0.3953.∎
\[ \Lambda_{\mathrm{off}} \in \Big[A_{\mathrm{off}} \mp \tfrac{T}{N}\Big], \quad T^2 = \frac{(r_1 - q_1 a_1^2)(q_1 - q_0)}{q_1 q_0} \]
Proposition 4
The sharp identified interval in the one-spike geometry, in closed form. At published one-minute statistics it is 7.4 to 8.9 times wider than the coefficient and contains zero: not identified even in sign.
\[ \psi_K(\widehat A) = \frac{\min_{D,\,\operatorname{rank}R \le K} \lVert \widehat A - D - R\rVert_F}{\lVert \widehat A - \operatorname{diag}\widehat A\rVert_F} \]
Diagnostic
Zero under pure confounding, so a materially positive value refutes it. Off-diagonal magnitude carries no information; only departure from \(\mathcal{D}_K\) does.
\[ C(x, A) - C(x, \Lambda) = x^{\top} G x \;\overset{\text{one-spike}}{=}\; g\,(\mathbf 1^{\top} x)^2 \]
Theorem 6
The execution-cost error is a low-rank quadratic form. It vanishes on a subspace of dimension at least \(N - K - \operatorname{rank}(B)\), 27 in the fixture, so a dollar-neutral trade has a point-identified cost even though the matrix behind it does not.
\[ x^{*}(\pi) = \tfrac12 \lambda\, M(\pi)^{-1} c, \qquad M(\pi) = A_s + \pi\,\mathbf 1 \mathbf 1^{\top} \]
Execution
The minimax-cost schedule under \(c^{\top}x = q\), at \(\pi = T/N\). It matches a 20,000-point grid search exactly and improves the worst case by 3.11% for a general target.
Figure 1. Identified intervals for an off-diagonal entry at the source-matched calibration (\(N = 30\)).
Figure 2. \(\psi_K\) across assumed factor counts: the elbow reads the true \(K = 3\) off the coefficient matrix.
Figure 3. Execution-cost error as a share of true cost, one-spike geometry.

Connections: execution, quadratic covariation and Kyle

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

Almgren-Chriss execution

\[ x(t) = X\,\frac{\sinh\big(\kappa(T - t)\big)}{\sinh(\kappa T)}, \qquad \kappa = \sqrt{\lambda\sigma^2/\eta} \]

Liquidating \(X\) shares with temporary impact \(\eta\) under a mean-variance objective with risk aversion \(\lambda\) gives this schedule. With many assets, \(\kappa\) becomes a matrix built from the covariance and the impact matrix, so the schedule depends on \(\Lambda\) exactly where identification fails. The figure below shows what a misestimated impact does to the plan.

Sampling faster cannot help

\[ \widehat\Lambda = \frac{d\langle r, q\rangle_t}{d\langle q, q\rangle_t}, \qquad \langle r, q\rangle_t = \big(\Lambda\Sigma_{qq} + \Gamma\Sigma_f\Delta_f^{\top}\big)\,t \]

In continuous time the regression slope is a ratio of quadratic covariations, which high-frequency data recovers exactly. But Proposition 3 says those second moments do not determine \(\Lambda\): the confounding gap \(\Gamma\Sigma_f\Delta_f^{\top}\Sigma_{qq}^{-1}\) survives any sampling frequency. Partial identification is a property of the model, not of the sample size.

Where real cross-impact would come from

\[ \lambda_{\text{Kyle}} = \frac{\sqrt{\Sigma_0}}{2\,\sigma_u} \]

In Kyle's model an informed trader and noise traders produce linear price impact. Correlated private information across assets is a structural reason \(\Lambda\) could be genuinely non-diagonal, and it is precisely the story the data cannot tell apart from a common factor.

Optimal schedulePlanned with misestimated impactTrade evenly

Illustration. Almgren-Chriss holdings over the trading window, with unit volatility and impact. Overestimating impact makes the plan trade too slowly and carry more risk; underestimating it makes it trade too fast and pay more impact. The default error is the 54.23% from Figure 3, used only as an illustrative size.

Verification and build

The derivation was frozen before any simulation code existed (gate G0). The known-truth experiment ran on one master draw split into 100 immutable shards of 100,000 observations, each publishing only mergeable sufficient statistics: all 1,800 coefficient targets fall inside their simultaneous intervals, the maximum relative discrepancy is 0.00056 against a preregistered 0.001, and the replay reproduces the summary byte for byte (gate G1).

Gate G0 freezes the mathematics before code; gate G1 is the known-truth verification.
Figure 4. Calibrating the \(\psi_K\) test by parametric bootstrap. Left: size at nominal 0.05, valid only above roughly \(T = 5N^2\). Right: power at \(T = 5{,}000\).
uv sync --locked --extra dev
make check     # ruff lint and format, strict mypy, 410 tests, smoke, drift
make exhibits  # regenerate every manuscript number; fails on drift
make paper     # build the preprint PDF

Econometrics, partial identification, factor models, market microstructure, parametric bootstrap, sharded simulation, strict typing, Python.

The hidden options market

FLEX inventory in clearing data, and whether it predicts the visible market

August to September 2026.

0activity dates
0underlyings
0series-days
0contracts open
0mark value
From a free daily clearing report to a tested panel. The black boxes are where the two measurement errors below are corrected.

The story. FLEX options are negotiated bilaterally, cleared centrally, and absent from the listed option chains most option research uses. The Options Clearing Corporation still publishes every cleared FLEX series, daily, for free. I built the pipeline that turns those reports into a panel, found two ways the obvious reading of the data is wrong, and then asked whether this hidden inventory predicts the visible market. It does not.

Figure 1. FLEX open interest and mark value on the first and last dates of the panel.

Errata: what the naive reading gets wrong

Naive readingWhat the data says
An off-grid strike marks a FLEX seriesOver 486,895 cleared series that rule has precision 0.67, recall 0.34: two thirds of FLEX is struck on the grid.
A cleared series is a positionCounted leg by leg, a butterfly is credited with 49x its maximum value.
Four legs, four exposuresA four-leg SPY collar is counted about 3x over.
Strike notional measures exposureA synthetic long call struck at $1.87 against a $765 spot is recorded as $7m when its exposure is $2.9bn.
File each report by the date it was fetchedOCC publishes overnight, so a capture taken during a session would enter one settlement twice; reports are filed by activity date.

Scoring the strike rule \(\hat F\) against the report's ground truth \(F\): \(\;\text{precision} = |\hat F \cap F| / |\hat F| = 0.67,\ \ \text{recall} = |\hat F \cap F| / |F| = 0.34.\) I used that rule first, replaced it, and kept the superseded census public.

Each leg's valueThe butterfly

Live. Why legs must be grouped, with illustrative strikes: long one call at 90, short two at 100, long one at 110, valued at expiry. Each leg can be worth a great deal while the structure is never worth more than 10.

Does hidden inventory predict the visible market?

A dealer on the other side of a FLEX trade hedges in the listed market and carries the residual convexity. Inventory for day \(t\) is published overnight, so every predictor is measured at \(t\) or earlier and every outcome strictly after:

\[ y_{i,t+1} = \alpha_i + \delta_t + \beta\,\mathrm{flex}_{i,t} + \gamma^{\top} z_{i,t} + \varepsilon_{i,t+1}, \tag{1} \]

with underlying and date fixed effects, errors clustered two ways, and controls for trailing realized volatility at 5 and 22 days, same-day absolute return and log dollar volume, over 58,014 underlying-days.

Genuine estimatesPlacebo

Four inventory measures, each estimated with equation (1) on next-day absolute returns.

The headline estimate is −0.00030 with t = −1.37. On an earlier, smaller panel it reached t = −1.95, the kind of number that gets written up as suggestive.

The placebo uses inventory dated after the outcome, so it cannot carry information. It lands on −0.00032, t = −1.37: whatever links inventory to volatility runs equally well backwards in time.

That is the signature of a common factor moving both, not of information flowing from one to the other.

A null is informative only in proportion to its power: \(\mathrm{MDE}_{80} = (z_{0.975} + z_{0.80})\,\widehat{\mathrm{SE}}(\hat\beta) \approx 2.80\,\widehat{\mathrm{SE}} \approx 0.0006\), against a typical daily absolute return near 0.02, so relative effects of 3 to 5% would have been detected.

Figure 2. The \(t\)-statistic of inventory innovation as the outcome horizon grows. Shaded: \(|t| < 1.96\).
Inventory innovationPlacebo
Figure 3. The same estimates from three panel start dates. The exact placebo match is specific to the first window; the absence of predictability is not.

Connections: Black-Scholes, gamma and densities

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

What a dealer carries

\[ d\Pi_t = \tfrac12\,\Gamma_t S_t^2\big(\sigma_{\text{realized}}^2 - \sigma_{\text{implied}}^2\big)\,dt, \qquad \Gamma = \frac{\varphi(d_1)}{S\,\sigma\sqrt{T}} \]

A delta-hedged option earns its gamma times the gap between realized and implied variance. A dealer on the other side of a FLEX trade hedges \(\Delta = \Phi(d_1)\) and keeps the convexity; short gamma, it rebalances with the move and can amplify volatility, long gamma it dampens it. That hedging feedback is the channel this study tests, and the placebo-tested null bounds it at the daily, per-underlying level.

A butterfly is a second derivative

\[ \lim_{h \to 0}\frac{C(K - h) - 2C(K) + C(K + h)}{h^2} = \frac{\partial^2 C}{\partial K^2} = e^{-rT} q(K) \]

By Breeden and Litzenberger, a butterfly with wing width \(h\), scaled by \(1/h^2\), prices the risk-neutral density \(q\) of the underlying at expiry. Its value lives in the combination of legs, never in a single leg, which is the deeper reason a cleared series cannot be read as a position. The figure below converges as you narrow the wings.

Risk-neutral densityButterfly price ÷ h²

Illustration. Black-Scholes call prices with \(S_0 = 100\), \(r = 0.02\), \(\sigma = 0.25\), \(T = 1\). Scaled butterflies across every strike converge to the lognormal density of \(S_T\) as the wings narrow.

uv venv && uv pip install -e '.[dev]'
.venv/bin/pytest
.venv/bin/python scripts/backfill_occ_flex.py --start 2025-07-23 --end 2026-08-17
.venv/bin/python scripts/run_predictive.py --start 2025-09-09 --end 2026-08-17

Data engineering, panel econometrics, fixed effects, two-way clustering, placebo tests, power analysis, options. 256 tests; code public, vendor captures kept in a private repository.

Nothing beats zero

Simulation calibration of forecast comparison on correlated equity panels

Preprint.

StackPython and Rust
Tests606
Replayhash-verified
Commits111

Abstract. Short-horizon equity forecasting papers report a model that beats a benchmark. This one audits the apparatus instead: every standard test is run on panels where the truth is fixed by construction, and its measured size is set beside its nominal size. Three components fail badly. The corrected apparatus holds its size, and on Borsa Istanbul data it returns a negative result with an explicit statement of what it could and could not have found. Open a row to see why each test fails.

Diebold-Mariano on panel rows, 4 correlated stocks0.2284, 4.6 times nominal

Stocks on the same session share market shocks, so for \(k\) units with average correlation \(\bar\rho\) the variance of a cross-sectional mean is inflated by \(\mathrm{VIF} = 1 + (k-1)\bar\rho\), and \(n_{\mathrm{eff}} = n/\mathrm{VIF}\). A test on rows overstates its statistic by \(\sqrt{\mathrm{VIF}}\):

\[ \alpha_{\mathrm{real}}(k) = 2\,\Phi\!\left(-\frac{z_{1-\alpha/2}}{\sqrt{1 + (k-1)\,\bar\rho}}\right). \]

At \(\bar\rho = 0.5697\) the simulation tracks this closed form to within 0.0064 from 0.11 to 0.79.∎

Diebold-Mariano on panel rows, 100 correlated stocks0.7913, 15.8 times nominal

Widening the panel makes it worse, because a session of \(k\) equicorrelated units carries only

\[ m(k) = \frac{k}{1 + (k-1)\,\bar\rho} \;\uparrow\; \frac{1}{\bar\rho} \]

independent observations (Proposition 1). Past a few dozen names, breadth buys almost no precision; only more sessions help. On the committed panel, 480 rows over 120 sessions carry about 177 observations.

Holm across a family of six, 30 stocks, rows0.5875, 11.8 times nominal

Holm controls the family-wise error rate only if each test in the family holds its own size. Fed row-level statistics, it inherits their inflation: a correction cannot repair the tests it is correcting.

Diebold-Mariano against a zero forecast0.9219, 18.4 times nominal

Against a zero benchmark the models are nested, and the squared-error differential is

\[ \mathbb{E}\big[e_{1,t}^2 - e_{0,t}^2\big] = \mathbb{E}\big[\hat y_t^{\,2}\big] - 2\,\mathbb{E}\big[\hat y_t\, y_t\big], \]

so a forecast with more variance is punished regardless of its covariance with the target. Clark-West subtracts the noise term, \(\hat f_t = e_{0,t}^2 - [e_{1,t}^2 - (\hat y_{0,t} - \hat y_{1,t})^2]\), and holds its size.

False Strategy Theorem threshold, any trial correlationcleared about half the time

The deflated Sharpe ratio compares the best of \(N\) trials with the expected maximum of \(N\) skill-free ones,

\[ SR^{*}_0 = \sqrt{V[\widehat{SR}]}\,\Big[(1-\gamma)\,\Phi^{-1}\!\big(1 - \tfrac1N\big) + \gamma\,\Phi^{-1}\!\big(1 - \tfrac1{Ne}\big)\Big], \]

but an expectation is not a critical value: a skill-free grid of 72 configurations clears it about half the time at every level of dependence. A joint stationary bootstrap gives a real 95th percentile, and the grid behaves like 3.77 independent trials, not 72.

Corrected apparatus: session aggregation, Clark-West, joint bootstrap0.0385 to 0.0565, holds

Every loss differential is averaged across the stocks present on a date, the nested comparison uses Clark-West, and the search is resampled jointly. Measured size stays between 0.0385 and 0.0565 at nominal 0.05.

Each bar runs from the nominal 5% (black tick) to the measured size.

Standard test, closed formCorrected apparatus

A test at 5% should reject a true null 5% of the time. On a single stock, Diebold-Mariano does. Keep scrolling to add stocks.

Four stocks correlated at 0.5697: the simulation measures 22.84%. The rows of a panel are not independent observations.

A hundred stocks: 79.13%. The curve is the closed form in the first row of the audit.

The corrected apparatus holds its size between 0.0385 and 0.0565.

Re-run the audit yourself. A Monte Carlo running in your browser: 120 sessions of \(k\) null loss differentials with \(\bar\rho = 0.5697\) per replication, tested on rows and on session averages. The row-level rate converges to the closed form, the session-level rate to 5%.

Figure 1. The information ceiling \(m(k)\) at \(\bar\rho = 0.5697\): four stocks on one session are worth about 1.48 independent observations, and no number of them is worth more than 1.76.

Borsa Istanbul, and what the design could detect

On a leakage-controlled panel of GARAN, ISCTR, KCHOL and THYAO with a target the backtest can trade, no model carries predictive content: White's Reality Check gives \(p = 0.9970\), Hansen's SPA \(p = 0.6891\). Inverting the test at size \(\alpha\) and power \(1-\beta\),

\[ \delta_{\min} = \big(t_{1-\alpha/2,\,n-1} + t_{1-\beta,\,n-1}\big)\cdot\mathrm{SE}(\bar d), \]

bounds what it could have seen: an out-of-sample \(R^2\) of 0.113 in closed form and 0.2012 by simulation, against the 0.01 this literature treats as meaningful.

Figure 3. The smallest detectable out-of-sample \(R^2\) against the effect that would matter, log scale.

Connections: why nothing beats zero is a Brownian-motion fact

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

Drift is invisible, volatility is not

\[ dX_t = \mu\,dt + \sigma\,dW_t: \qquad \operatorname{SE}(\hat\mu) = \frac{\sigma}{\sqrt T}, \qquad \frac{1}{T}\sum_k (\Delta X_k)^2 \xrightarrow{\;\Delta t \to 0\;} \sigma^2 \]

Sampling faster recovers \(\sigma\) exactly through quadratic variation, but the error in the drift depends only on calendar time. A return forecast is a statement about drift, so a short panel cannot resolve a plausible effect, however finely it is sampled: the same conclusion as the detectability bound, from first principles. The figure below runs it.

The maximum of noise

\[ \mathbb{E}\Big[\max_{i \le N} Z_i\Big] \sim \sqrt{2\log N}, \qquad \frac{\max_{i \le N} Z_i - b_N}{a_N} \Rightarrow \text{Gumbel} \]

The best of \(N\) skill-free strategies looks good by extreme-value theory alone. The deflated Sharpe threshold (6) refines this expectation; correlation between trials shrinks the effective \(N\), which is why 72 configurations behave like 3.77.

The Sharpe ratio of a Brownian strategy

\[ \widehat{SR} = \frac{\hat\mu}{\hat\sigma}, \qquad \operatorname{Var}\big(\widehat{SR}\big) \approx \frac{1 + SR^2/2}{n} \]

Because \(\hat\mu\) carries almost all the error, so does the Sharpe ratio. This variance is the \(V_{\mathrm{ind}}\) behind the effective-trial count in Proposition 3.

Simulation. Three hundred five-year histories of a stock with 6% annual drift and 20% volatility, each estimated from data at the chosen frequency. The volatility estimates collapse onto the truth as sampling gets finer; the drift estimates do not move, because their error depends on the five years, not on the number of observations.

Calibrate first, then measure. The black box is the known-truth generator every estimator is run against.
uv sync
make reproduce-committed        # replay the committed run, verify every artifact hash, offline
uv run bist-predict calibrate   # rebuild the calibration, about four minutes
make verify-claims RUN_ID=20260810T230615Z-d346031-2a71b8

Monte Carlo calibration, Diebold-Mariano, Clark-West, Holm, White's Reality Check, Hansen's SPA, deflated Sharpe ratio, stationary bootstrap. Python with a Rust indicator library, 606 tests, an external review and my response in the repository.

Crypto returns and the martingale null

Nested MSPE inference, purged walk-forward validation and multiple-testing control, with bootstrap bounds on net Sharpe ratios

Daily data, 23 July 2020 to 17 July 2026.

Every out-of-sample block of the study, in order: 35 folds of 63 days for BTC and ETH, 28 for SOL. Each is forecast by a model trained only on what came before it, minus a purge and a 5-day embargo. Hover a block.

Abstract. I test whether daily BTC, ETH and SOL returns can be forecast out of sample well enough to survive trading costs. Three assets, two horizons and six forecasters give 18 machine-learning settings against a martingale null. The headline depends entirely on the test, and the funnel below shows each one applied in turn: a faint and fragile statistical signal in BTC at one day, and no tradeable one anywhere.

The validation that cannot see the future

A \(h\)-step target resolves at \(t + h\), so the \(h\) rows before each test block are purged and a further embargo of \(e\) bars is imposed; every fold satisfies \(\max(\mathcal{T}_k) + h + e < \min(\mathcal{V}_k)\). Scalers are fitted inside each fold, and gradient boosting purges its own early-stopping slice, which a test asserts by corrupting exactly those rows.

Why Diebold-Mariano is the wrong test

Against the recursively estimated mean \(\hat y^{\,b}\), Diebold-Mariano compares squared errors with a Newey-West variance, but the random walk is nested in every model, so estimation noise reads as evidence for the benchmark. Clark-West removes it:

\[ \hat f_t = \big(y_t - \hat y^{\,b}_t\big)^2 - \Big[\big(y_t - \hat y^{\,m}_t\big)^2 - \big(\hat y^{\,b}_t - \hat y^{\,m}_t\big)^2\Big], \qquad \mathrm{CW} = \frac{\bar f}{\sqrt{\hat V_f / n}}. \]

Live. The target is pure noise and the larger model forecasts pure noise too, so the null is true by construction. Diebold-Mariano still convicts the model far above its nominal rate; Clark-West holds its size. This mirrors a test in the repository.

Multiple testing

With \(m = 18\) ordered p-values, Holm rejects step-down while \(p_{(i)} \le \alpha/(m - i + 1)\), and Benjamini-Hochberg step-up at the largest \(i\) with \(p_{(i)} \le i\alpha/m\).

HolmBenjamini-HochbergSurvives BHDoes not
Figure 1. The five smallest Clark-West p-values against both thresholds, log scale. BH is step-up: rank 2 falls under its line and carries rank 1 with it. Nothing falls under Holm's.
Figure 2. Net Sharpe ratios of the two survivors, with stationary-bootstrap 95% intervals. Their out-of-sample \(R^2\) is −0.0079 and +0.0031, and every high-Sharpe result in the study is long-only exposure in disguise.

Connections: martingales, regularization and scaling

Not part of the repository. How the same problem reads through stochastic calculus, pricing theory and learning theory.

The null is a martingale

\[ \mathbb{E}\big[r_{t+1} \mid \mathcal F_t\big] = 0 \iff \log P_t \text{ is a martingale} \]

Under the null, log prices are martingales; under a risk-neutral measure every discounted price is one, which is the Girsanov change of measure from the notebook. A forecasting edge is a predictable piece of drift, and drift is the quantity Brownian motion hides best.

Regularization interpolates to the benchmark

\[ \hat\beta_\lambda = \big(X^{\top}X + \lambda I\big)^{-1}X^{\top}y \;\xrightarrow{\;\lambda \to \infty\;}\; 0 \]

Ridge with infinite penalty is the zero forecast, the random walk itself; the elastic net shares that path, and gradient boosting \(F_m = F_{m-1} + \nu h_m\) with no trees is a constant. The benchmark sits inside every model family here, which is the nesting that biases Diebold-Mariano and the reason Clark-West is needed.

Variance grows linearly

\[ \operatorname{VR}(q) = \frac{\operatorname{Var}\big(r_t^{(q)}\big)}{q\,\operatorname{Var}(r_t)} = 1 \quad \text{under a random walk} \]

Brownian scaling, \(\operatorname{Var}(X_{t+q} - X_t) = q\sigma^2\), gives the Lo-MacKinlay variance-ratio test: a classic attack on the same null from the scaling side rather than the forecasting side.

make setup      # install the package and dev extras
make backtest   # download data, run the study, regenerate reports and figures, about 15 s
make test       # 154 tests, including the no-lookahead and purge guarantees
make check      # ruff, black, mypy, 80% coverage floor: the CI gate on Python 3.12 and 3.13

Time-series cross-validation, Clark-West, Diebold-Mariano, Holm, Benjamini-Hochberg, stationary bootstrap, deflated Sharpe ratio, scikit-learn. 154 tests at 96% coverage; the report is generated from the results table, never hand-edited.

Probability, in the margins

Eight theorems about random walks and Brownian motion, each simulated in your browser until it happens in front of you

Built for this site.

Why this page exists. Quantitative finance runs on a small number of deep facts about randomness. Random walks become Brownian motion under the right scaling; fair games produce long, lopsided streaks; a fair process strays exactly as far as the law of the iterated logarithm allows; no stopping rule beats a martingale. In continuous time, Brownian paths have quadratic variation equal to elapsed time, so the chain rule gains a term; integrals depend on where you evaluate them; changing the measure reweights paths without moving them; and prices are expectations that solve partial differential equations. Each plate below is one of those facts, computed exactly enough to watch. The same processes run through the rest of the site: the line under the top bar is a Brownian path whose quadratic variation is your reading progress, and the rules between sections are Brownian bridges.

Part one. Discrete time

Plate I

Donsker: a coin becomes Brownian motion

Rescale the first \(n\) steps of a random walk by \(1/\sqrt n\). As \(n\) grows, the law of the whole path converges to Wiener measure, and it does not matter what the steps were, only that they have mean zero and variance one. This is why binomial trees converge to Black-Scholes.

\[ W^{(n)}_t = \frac{S_{\lfloor nt\rfloor}}{\sqrt n} \;\Longrightarrow\; W \quad \text{in } D[0,1] \]
Plate II

The arcsine law: fair games are lopsided

Play a fair game for a long time and record the fraction of time you are ahead. The most likely outcomes are almost always ahead or almost always behind; an even split is the least likely. It is why a long winning streak in a track record is weak evidence of skill.

\[ \mathbb{P}\big(\Gamma_n \le x\big) \;\longrightarrow\; \frac{2}{\pi}\arcsin\sqrt{x} \]
Plate III

The law of the iterated logarithm

The central limit theorem says \(S_n/\sqrt n\) is roughly normal at any fixed time. Khinchin's law says how far it strays along the whole path: out to \(\sqrt{2\log\log n}\), touched infinitely often and exceeded only finitely often. It is the boundary behind sequential testing.

\[ \limsup_{n\to\infty} \frac{S_n}{\sqrt{2n\log\log n}} = 1 \quad \text{a.s.} \]
Plate IV

Optional stopping: no system beats a fair game

A gambler starts with \(x\) and plays a fair coin until ruin at 0 or a target \(N\). Because \(S_n\) and \(S_n^2 - n\) are martingales and the stopping time is integrable, the probability of reaching the target and the expected length of the game follow without any computation of paths.

\[ \begin{aligned} \mathbb{E}[S_\tau] = x &\;\Rightarrow\; \mathbb{P}(S_\tau = N) = \tfrac{x}{N}, \\ \mathbb{E}[S_\tau^2 - \tau] = x^2 &\;\Rightarrow\; \mathbb{E}\,\tau = x(N - x) \end{aligned} \]

Part two. Continuous time

Plate V

Quadratic variation

Cut a Brownian path at the dyadic times \(k2^{-n}\). The squared increments sum to elapsed time while the path's length diverges like \(2^{n/2}\). Every smooth function has zero quadratic variation, so this single number is what separates stochastic calculus from the ordinary kind, and why realized variance estimates volatility.

\[ \begin{aligned} &\textstyle\sum_k \big(W_{t_{k+1}} - W_{t_k}\big)^2 \xrightarrow{\text{a.s.}} t, \\ &\textstyle\mathbb{E}\sum_k \big|\Delta W\big| = \sqrt{2/\pi}\;2^{n/2} \end{aligned} \]
Plate VI

Itô against Stratonovich

Because the increments do not vanish fast enough, a Riemann sum for \(\int W\,dW\) depends on where each interval is evaluated: the left endpoint gives Itô, the midpoint gives Stratonovich, and they differ by half the quadratic variation. Finance uses Itô because a strategy chooses its position before the price moves.

\[ \int_0^t W\,dW = \tfrac12 W_t^2 - \tfrac12 t, \qquad \int_0^t W \circ dW = \tfrac12 W_t^2 \]
Left-point sumsMidpoint sums
Plate VII

Girsanov: change the measure, not the paths

Weight each path by the likelihood ratio and \(W_t - \theta t\) becomes a Brownian motion: the same paths now carry drift \(\theta\). Risk-neutral pricing is this reweighting. The effective sample size shows the price: tilt far and a handful of paths carry all the weight, which is why importance sampling degenerates.

\[ \frac{dQ}{dP}\bigg|_{\mathcal F_T} = \exp\!\Big(\theta W_T - \tfrac12\theta^2 T\Big) \]
Plate VIII

Black-Scholes as a Feynman-Kac expectation

Under the risk-neutral measure, \(dS_t = rS_t\,dt + \sigma S_t\,dW_t\). By Feynman-Kac, the price of a claim on \(S_T\) is both an expectation over paths and the solution of a partial differential equation; for a call, the closed form below. Monte Carlo reaches it at the rate \(1/\sqrt n\).

\[ \begin{aligned} &\partial_t u + r x\,\partial_x u + \tfrac12\sigma^2 x^2\,\partial_{xx} u = r u, \\ &C = S_0\,\Phi(d_1) - K e^{-rT}\,\Phi(d_2) \end{aligned} \]

Legal and privacy

Last updated 2 October 2026.