Journal
Work in the open
The phone-sized model is being built with every decision pre-registered and every result journaled, corrections included. Entries appear here as they are reviewed. Below, the research list, and the compression preprint once it is ready.
Research, with the negative results kept
- Local inference of a 284-billion-parameter mixture-of-experts model on one workstation: decode raised from 7.3 to about 20 tokens per second; 190+ numbered journal entries and pre-registered decisions.
- Lossless LLM-driven text compression: 0.882 bits per byte on enwik8 against 7-Zip's 1.989, with a measured ratio-versus-speed frontier. Preprint in preparation.
- A deterministic context manager for long agent sessions: byte-exact log, lean resident window, 93-100% recall of exact facts at 90-99% token reduction, claim narrowed after an adversarial audit.
- On-device model selection for phones: an MMLU-Pro harness projected to memory bandwidth across 298 candidate models.
- Shepherd: a self-hosted endpoint control plane and Windows agent with tenant isolation, signed tasking, and a hash-chained audit log, built so an AI agent can operate a fleet safely. Working prototype, ten decision records.
- A de-identified calibration corpus built from real platform traffic, with two adversarial privacy audits; the first said "do not ship."