Files
pyRBM/src/tests/README.md
T
2026-05-30 20:44:13 +02:00

7.6 KiB
Raw Blame History

Tests

Run any test from the src/tests/ directory:

cd src/tests
python test_<name>.py

Self-contained tests

These generate their own data and run without external files.

File RBM type What it tests
test_binary_autoencoder.py BB-RBM Auto-encoder on synthetic bars-and-stripes binary images (8×8). Trains, reconstructs, and plots original vs. reconstruction.
test_model.py BB-RBM (4 layers) Deep auto-encoder on random binary data (1024-dim). Prints per-pattern reconstruction error.
test_xor.py BB-RBM Auto-encoder on XOR-like 3-bit patterns using the older Layer API.
test_gaussian_autoencoder.py GB-RBM Auto-encoder on synthetic Gaussian blob images (8×8). Trains, reconstructs, and plots original vs. reconstruction.
test_linear.py GB-RBM or GG-RBM GB/GG-RBM on random continuous data (3000-dim). Toggle do_gaussian_hidden at the top of the file.

Tests requiring external image data

These load images from a fixed path on disk.

File RBM type What it tests
test_faces_sub_image.py GB-RBM Auto-encoder on 32×32 patches from Caltech WebFaces (/media/jens/cifs/bilder/MachineVision/Caltech_WebFaces/). Uses SubImage with 50% overlap (stride=16); reconstruction blends overlapping patches by pixel-averaging. Supports RGB and grayscale modes. CLI args: --grayscale, --load_model, --do_train, --l1_lambda. After training: shows learned weight filters, patch-level reconstructions, and a full overlap-blended image reconstruction. Default: 256 hidden units, 1000 epochs, mini_batch=1000. L1=0.01 is the recommended sweet spot — produces structured edge/texture filters with 15% sparsity while keeping MSE0.242. L1=0.05 collapses most filters to blob detectors (58% sparsity, MSE0.261).
test_sub_image.py Unit test for SubImage (patch extraction). No RBM involved.
test_conv2d.py Unit test for conv2d. No RBM involved.

Tests requiring external data

These load data from /home/jens/work/repos/Rbm/ (Armadillo .dat files).

File RBM type What it tests
test_norbs.py GB-RBM or GG-RBM Single-layer auto-encoder on NORB small (96×96 images). Toggle DO_HIDDEN_GAUSSIAN at the top of the file.
test_deep_norbs.py GB-RBM + BB-RBM Two-layer deep auto-encoder: loads pre-trained GB-RBM weights, trains a BB-RBM on top.
test_learn_norbs_labels.py GB-RBM + BB-RBM Multimodal model combining NORB image features with one-hot label embeddings in a joint hidden layer.
test_rbm.py Any (from .prj file) Full pipeline via StackFactory. Takes a project name as CLI argument (python test_rbm.py <name>). Uses OpenCV for interactive reconstruction display.

Tests requiring saved state

These expect a trained model to already exist in results/.

File Prerequisite
test_label.py Requires a saved Label fitter. Trains a binary label encoder (vocabulary of 100 integers → 5×5 binary codes) and visualises encoded labels. Set do_train = True on first run.
test_learn_encoded_labels.py Requires a saved one-hot label fitter. Trains a BB-RBM on one-hot encoded integer labels and plots hidden representation, reconstruction, and original. Set do_train = True on first run.

Implementation notes

Bug fixes applied (see rbm/train.py, rbm/matrix.py, rbm/label.py)

train.py — momentum reset every epoch (critical)

grad_zero() was called inside the epoch loop, zeroing the momentum accumulator before every gradient update. For the default full-batch case (one mini-batch per epoch), momentum had zero effect — every step was just lr * dw. Moved grad_zero() outside the loop so momentum accumulates correctly across epochs. GB-RBM reconstruction MSE dropped from ~0.22 to ~0.014 on the Gaussian blob test.

train.pycd_gaussian_binary negative phase (two fixes)

  1. Spurious Gaussian noiseh_probs_neg was computed from data_neg + N(0,1) instead of data_neg. Noise injected variance into every gradient step.
  2. Soft hidden states for reconstructiondata_neg was computed from h_probs_pos (probabilities) instead of sampled binary states, biasing the fantasy particle toward the mean.

train.pycd_gaussian_gaussian negative phase (same pattern)

Same two bugs as cd_gaussian_binary: noise was added to data_neg before computing h_probs_neg, and h_probs_pos (mean) was used to drive the negative reconstruction instead of a Gaussian sample. Fixed to use sampled h for reconstruction and clean means for weight updates.

matrix.pyrms_error wrong denominator

d_err_squared[1] (row 1 of the matrix) was used instead of d_err_squared.shape[1] (column count). Dead code but now correct.

label.py — CuPy compatibility and incomplete implementation

  • _dec_binary used Python's reversed() on a CuPy array (unsupported) — replaced with np.flip().
  • label2vec_onehot returned a Python list instead of a stacked array — now uses np.stack().
  • vec2label_onehot was unimplemented (returned None) — now uses np.argmax().

entity.py — L1 and L2 regularisation gradients wrong

grad_compute computed scalar norms and multiplied by W, giving gradients that scaled with the total weight magnitude rather than the correct per-element derivatives:

  • L1: gradient should be λ·sign(W), not λ·‖W‖₁·W
  • L2: gradient should be 2λ·W, not λ·‖W‖₂²·W

The combined (l1+l2)·W term also incorrectly mixed both penalties into one. Fixed to apply each gradient independently. weight_decay (the correct λW form of L2) is kept as a separate term.

image.pynormalize() fragile and verbose

The repeat/reshape approach was replaced with keepdims=True broadcasting and a std < 1e-8 guard to prevent NaN on constant patches. The duplicate normalization logic in test_faces_sub_image.py:load_image_patches was removed in favour of calling normalize() directly.

train.py — no data shuffling between epochs

Mini-batches were always sliced in the same order every epoch, introducing systematic gradient bias for mini-batch training. A random permutation is now applied to the full batch at the start of each epoch.

train.pyif chains instead of elif/else for CD function selection

All four if statements were evaluated even after a match, and an unrecognised entity type would leave cd_func = None, crashing later with a cryptic TypeError. Replaced with elif/else that raises ValueError immediately.

entity.py — L1/L2 regularisation divided by mini-batch size

state_adjust(grad, k=1/N) applied the 1/N scale to the entire gradient, including the regularisation terms. The CD gradient is a batch sum so 1/N normalisation is correct; L1/L2/weight-decay penalties are per-weight and batch-size independent. Dividing them by N made the effective regularisation strength lambda / N, so e.g. l1_lambda=0.5 with mini_batch_size=1000 was equivalent to an effective L1 of 0.0005 — too small to have any visible effect on filters. Fixed by moving the regularisation update into state_adjust where it is applied at full strength before the 1/N CD update.


RBM type reference

Type Visible Hidden EntityParams
BB-RBM Binary Binary EntityParams()
GB-RBM Gaussian Binary EntityParams(do_gaussian_visible=True)
BG-RBM Binary Gaussian EntityParams(do_gaussian_hidden=True)
GG-RBM Gaussian Gaussian EntityParams(do_gaussian_visible=True, do_gaussian_hidden=True)