Ideas explored, worked out and written up
My own axiomatic and computational framework: build the hyperreals ℝ* algebraically (like ℂ = ℝ(i)) from a single gauging law ε·ω = 1, with a full executable Lean 4 model. Used to derive an algebraic derivative, integral, and Dirac delta, and to replace Kolmogorov's σ-algebra + measure with an algebraic, codimension-indexed atom mass — resolving “probability zero but not impossible” with a genuine order of infinitesimal-smallness classes instead of one flat zero. Scope is stated honestly: expectation, conditional probability across mixed orders, and LLN/CLT are open gaps, not glossed over.
Read PDF CodePrediction-error-driven Boltzmann exploration temperature — my own method and a multi-week experimental program, from tabular DQN and GRPO up to a per-state variant on continuous control. Beats VDBE-Softmax on classic control; a PPO+GAE+LSTM version comes close to solving BipedalWalker-v3 (peak 289.1, short of the standard 300 threshold), backed by a component-ablation study. Claude's survey situates the general principle relative to VDBE-Softmax, UQL, and CBSQL.
Read PDF CodeByte-level, tokenizer-free transformers with a local/global hierarchy. Survey finds this is already realized at scale by MEGABYTE and Meta's Byte Latent Transformer. A small experiment trains a toy version on real Commodore 64 game binaries — a domain not covered elsewhere — with one clean result and one honestly-reported negative result.
Read PDF CodeA working harness — my own project, actively developed — that wraps a live Commodore 64 emulator as a reinforcement-learning environment and uses an LLM-driven agent to choose between hard-coded, tabular, and distilled neural policies per game. No published academic work targets C64 specifically. Real per-game results reported honestly, plus a new, tested improvement to the system's policy promotion gate.
Read PDF CodeDistilling a slow, expensive policy (e.g. one driven by a language model) into a fast network via DAgger, gated by a held-out promotion test. The core mechanism is established, but the practical cost of querying an expensive teacher every round is not. A confidence-margin filter cuts expert queries by over half with no loss in final policy quality, tested on a real distillation pipeline — part of the same Transplorer project above.
Read PDF Code