Research Papers

Ideas explored, worked out and written up

How these came to be: each of these ideas, started as my own prior exploration — Claude did literature survey against existing work, proposed and ran additional follow-on experiments, I directed the process Where the harness needed assistance. Claude wrote up the results, I reviewed the outcome. Feedback and corrections welcome — let me know if anything reads off.

Hyperreal Numbers: A Codimension-Indexed Reformulation of Probability main paper

My own axiomatic and computational framework: build the hyperreals ℝ* algebraically (like ℂ = ℝ(i)) from a single gauging law ε·ω = 1, with a full executable Lean 4 model. Used to derive an algebraic derivative, integral, and Dirac delta, and to replace Kolmogorov's σ-algebra + measure with an algebraic, codimension-indexed atom mass — resolving “probability zero but not impossible” with a genuine order of infinitesimal-smallness classes instead of one flat zero. Scope is stated honestly: expectation, conditional probability across mixed orders, and LLN/CLT are open gaps, not glossed over.

Read PDF Code

Error-Rate Temperature Steering (ERTS) real method, real results

Prediction-error-driven Boltzmann exploration temperature — my own method and a multi-week experimental program, from tabular DQN and GRPO up to a per-state variant on continuous control. Beats VDBE-Softmax on classic control; a PPO+GAE+LSTM version comes close to solving BipedalWalker-v3 (peak 289.1, short of the standard 300 threshold), backed by a component-ablation study. Claude's survey situates the general principle relative to VDBE-Softmax, UQL, and CBSQL.

Read PDF Code

Bytecoder: Tokenizer-Free Hierarchical Byte Models idea already covered

Byte-level, tokenizer-free transformers with a local/global hierarchy. Survey finds this is already realized at scale by MEGABYTE and Meta's Byte Latent Transformer. A small experiment trains a toy version on real Commodore 64 game binaries — a domain not covered elsewhere — with one clean result and one honestly-reported negative result.

Read PDF Code

Transplorer: An LLM-Orchestrated C64 Game-Solving Harness real system, real results

A working harness — my own project, actively developed — that wraps a live Commodore 64 emulator as a reinforcement-learning environment and uses an LLM-driven agent to choose between hard-coded, tabular, and distilled neural policies per game. No published academic work targets C64 specifically. Real per-game results reported honestly, plus a new, tested improvement to the system's policy promotion gate.

Read PDF Code

Query-Efficient DAgger for Distilling Expensive Teacher Policies open gap, tested

Distilling a slow, expensive policy (e.g. one driven by a language model) into a fast network via DAgger, gated by a held-out promotion test. The core mechanism is established, but the practical cost of querying an expensive teacher every round is not. A confidence-margin filter cuts expert queries by over half with no loss in final policy quality, tested on a real distillation pipeline — part of the same Transplorer project above.

Read PDF Code