Single-layer Deep stack (28x28->256), has saved weights and a larger
training batch than the other mnist variants (test run took noticeably
longer). Passes: reconstruction error 0.005 vs. mean-baseline 0.067.
9/12 pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016K8Gu7Qejd11JbdiHZqYAs