- Moved src/tests/ → tests/
- Added test_* functions with assertions to script-style test files
- Added main() to each so IDEs offer it as a separate run target from pytest
- Fixed cupy_test.py: remove spurious x_gpu += x_cpu, fix duplicate xlabel
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add PATCH/STRIDE params; STRIDE=4 for 50% overlap (49 patches per image)
- Reconstruct 32x32 by averaging overlapping patch contributions (assemble_overlapping pattern from test_faces_sub_image)
- All cells updated to use PATCH/STRIDE throughout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Tile the single sentence 500× to form a (500, T, 40) batch so matrix ops
are (500, 168) instead of (1, 168). Note: with identical rows the gradient
is mathematically equivalent to batch=1; real speedup requires diverse data.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
SENTENCE now accepts mixed case and strips non-vocab chars automatically.
Extended to ~200 chars using the Moby Dick opening passage.
Visualisation cells use adaptive tick spacing (max 40 labels) to handle
longer sequences without crowding.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remove context33.training.dat loading (and read_armadillo import).
Training data is now a hardcoded 20-character sentence encoded with the
same 40-char vocabulary as moby_rnn.ipynb (space .!? A-Z 0-9).
SENTENCE = "MOBY DICK IS A WHALE"
Results: 95% reconstruction, 95% next-step prediction (one miss each
on the first character where context is zero).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The training.dat file stores samples LIFO (last added in the C++ GUI = row 0).
Reversing the rows restores chronological order, matching the C++ sequence.
Next-step prediction remains 100%: model predicts x_{t+1} from h_t correctly
in the C++ insertion order 1 5 6 9 B F I L Q R U X.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Notebook reconstructs the C++ context33.prj configuration exactly:
- shared-weights StackRnn (1 entity, reused every time step)
- visible = [context(128) | x_t(40)] = 168 units, hidden = 128
- all hyperparams from .prj: lr=0.05, momentum=0.5, epochs=1000,
mini_batch=100, gibbs=3, rao_blackwell=True, weight_decay=0
Training data loaded directly from context33.training.dat (Armadillo format,
12 × 168): sensory one-hot chars decoded as 'XURQLIFB9651'.
Key design: training pairs are (h_t, x_{t+1}) not (h_{t-1}, x_t) —
context is first advanced by seeing x_t, then the model is trained to
predict the next character x_{t+1} from that context. Evaluation using
clamped Gibbs on h_t achieves 100% next-step prediction accuracy.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously a single shared-weight Entity processed every time step. Now:
- Shared mode (1 layer via make_layer): original behaviour unchanged
- Unrolled mode (N layers via make_unrolled): layers[t] owns W_t, b_v_t, b_h_t
New API:
make_unrolled(T, sensory_size, h_size, ...) → list[Layer]
next_entity() → Entity for the upcoming step() call
current_entity() → Entity from the most recent step() call
is_shared → bool
moby_rnn.ipynb: switch Build model cell to make_unrolled(T=100); update
predict_next() to use rnn.next_entity() for position-correct Gibbs sampling
README_moby_rnn.md: redraw temporal-unrolling ASCII art showing per-position
weights W_t; update parameter count and checkpoint file listing
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- README_moby_rnn.md: single-step and temporally-unrolled ASCII diagrams,
configuration table, notebook cell guide, generation API docs, references
- moby_rnn.ipynb: move all constants (T, NUM_SEQ, EVAL_CHARS, PRED_CHARS,
N_GIBBS, SEED, SEQ_IDX) into the Build model cell under a Configuration
header; rename H_SIZE → CONTEXT_SIZE throughout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Implements StackRnn: a cascaded/recurrent RBM where the visible layer at
each time step is the concatenation of the previous hidden state (context)
and the current sensory input — V[t] = [context | x_t]. Weights are
shared across time steps (RTRBM-style concatenation variant).
- stack_rnn.py: StackRnn with step(), reconstruct(), reset(), train(),
make_layer() factory; supports greedy layer-wise training over sequences
of shape (num_seq, T, sensory_size)
- test_rnn.py: single-layer, two-layer, and save/load tests
- moby_rnn.ipynb: character-level language model on Moby Dick; one-hot
encoding, clamped-Gibbs next-char prediction, free text generation,
hidden-state trace and character-distribution visualisations
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
test_learn_encoded_labels: add .get() for cupy→numpy conversion before imshow;
reduce num_epochs 10000→1000.
test_linear: uncomment model.load() to resume from saved weights.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
status.py: add CheckpointStatus(save_fn, update_interval) — prints progress
and calls save_fn at every report interval to persist model state mid-training.
model.py: Model.train() now accepts an optional Status instance; defaults to
plain Status() if none provided.
test_faces_sub_image.py: use CheckpointStatus(model.save) during training.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
train.py: add _to_gpu() helper; convert mini-batches from CPU numpy to GPU
CuPy on the fly so large datasets can stay in RAM. Use numpy permutation for
numpy batches during shuffle. Limit status err_rms to first 5000 samples to
avoid GPU OOM on status checks.
model.py: replace np.copy() with np.array() — cupy.copy() does not convert
numpy arrays to CuPy; cupy.array() does. Fixes entity.forward() crash after
training when batch was passed as numpy.
test_faces_sub_image.py: load_patches normalizes on CPU and returns numpy,
avoiding multi-copy GPU OOM for large datasets (500 images ≈ 1.4 GB).
show_reconstructions converts numpy patches to CuPy before inference.
Add --n_images CLI arg.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove cd_jens: dead code never called, contained multiple bugs.
- cd_gaussian_binary: dbv formula simplified — b_v terms cancelled between
positive and negative phases, making the subtraction redundant.
- cd_binary_gaussian: Gibbs loop (CD-k > 1) now properly samples both visible
(sample(prob(...))) and hidden (+ sample_gaussian) states when
do_rao_blackwell=False. Previously used means throughout, giving biased
gradient estimates for k > 1.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Shuffling eliminates systematic gradient bias from fixed mini-batch ordering.
elif/else raises ValueError for unrecognised entity types instead of silently
calling None and crashing with a cryptic TypeError.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Regularisation terms (L1, L2, weight_decay) were being divided by mini_batch_size
in state_adjust along with the CD gradient. The CD gradient is a batch sum so
1/N normalisation is correct; regularisation penalties are per-weight and
batch-size independent. Separating them makes l1_lambda/l2_lambda directly
interpretable regardless of batch size.
Also adds --l1_lambda CLI arg to test_faces_sub_image.py and sets num_epochs=1000.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previous code computed scalar norms and multiplied by W, giving gradients
that scaled with total weight magnitude rather than the correct per-element
derivatives:
L1: gradient is λ·sign(W), not λ·‖W‖₁·W
L2: gradient is 2λ·W, not λ·‖W‖₂²·W
Also removed the combined (l1+l2)*W term which incorrectly mixed the two.
weight_decay (correct L2 form λW) is kept as a separate term.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
image.py: replace repeat/reshape with keepdims=True; guard std < 1e-8 to
prevent NaN on constant patches (previously relied on call-site pre-filtering).
test_faces_sub_image.py: load_image_patches now calls normalize() directly
instead of duplicating the per-sample normalization logic.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
GRAYSCALE flag drives N_CH (1 or 3) and N_VIS (1024 or 3072) at runtime.
_img_to_chw() handles BGR→gray or BGR→RGB conversion.
All visualisation functions use cmap='gray' when GRAYSCALE is set.
Separate model name (faces_sub_image_gray) keeps weights isolated.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>