Files
brewpi/README.md
T
jensandClaude Sonnet 4.6 8c23a2c6ac Clear the Sud Load-rejection Error so it doesn't get replayed as stale
WsServerMultiUser.global_state persists every message ever broadcast
and replays the full accumulated state to any client that subscribes
- by design, so a client connecting fresh or reconnecting mid-run
always gets the same complete picture (see README's new "State replay
on connect" section). That mechanism has no concept of "transient":
the {'Sud': {'Error': ...}} added for Load-rejection got treated as
durable state just like everything else, so any client connecting
after the fact - even one that never touched Load - got the stale
error replayed on connect. Confirmed live before this fix.

tasks/sud.py: send {'Sud': {'Error': None}} immediately after the
error itself, clearing it back to neutral in global_state.

client/brewpi_gui.py: on_sud_changed() treats a falsy Error as a
no-op instead of popping an empty dialog.

README.md: documents the underlying principle (client state replay
on connect, and the corollary that one-shot events need explicit
clearing) under a new Architecture subsection, and refreshes the Sud
channel/forecast sections to match this session's other changes
(Elapsed-based dynamic forecast anchoring instead of wall-clock time
times the nominal warp factor, Load's Error reply).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LhiQe64F74uHV8jzuoSa5K
2026-06-22 13:20:24 +02:00

334 lines
18 KiB
Markdown

# BrewPi
A Python-based controller for automating the mash/brewing process of beer: it
holds a pot of liquid at target temperatures (or ramps it at a target heating
rate) according to a configurable mash schedule, drives a heater and stirrer,
and exposes live control/telemetry over a WebSocket so a desktop GUI (or any
other client) can monitor and steer the brew.
The project began as a simulation/control-theory playground (Smith-predictor
temperature control, pot transport-delay model, see
[`docs/NonLinMPC.pdf`](docs/NonLinMPC.pdf)) and has grown real-hardware
backends for an induction hob and an RTD temperature probe.
## Architecture
```
server/brewpi.py Server entry point: wires sensor, pot/plant,
heater, temperature controller and stirrer
together, runs them as asyncio tasks, and
serves state/commands over a WebSocket.
client/brewpi_gui.py PyQt5 desktop client (brewpi.ui) that connects
to the server's WebSocket, displays live
temperature/power/state, and lets the user set
target temperature, heat rate, stirrer speed,
ambient temperature, and switch the heater/
stirrer/temperature controller on or off (plus
a "reset Pot to ambient" button, shown only
when the server's plant is simulated) - plus
Sud control (New/Load/Save/Start/Pause/Stop)
and its forecast plot on the Automatic tab.
client/user_config.py Small JSON-backed key/value store
(~/.config/brewpi/gui.json) for GUI preferences
that should survive restarts (last Sud file
dialog directory, last ambient temperature) -
distinct from the server's config.json and
from sude/*.json schedules.
components/ Pluggable building blocks behind factories:
pid/ temperature controllers: plain PID
("Normal") or PID + Smith predictor
("Smith", runs two internal pot models —
one with the plant's transport delay, one
without — to compensate for the dead time).
Both are a 4-state FSM (IDLE/HEAT/HOLD/COOL)
with a master enabled switch - IDLE means
disabled (output forced to 0); COOL handles
a target below the current temperature with
its own gains, output forced to 0 only by
the actuator that can't act on it (e.g. a
heat-only heater), not by the controller
assuming it can't cool
plant/ pot thermal model (transport-delay line)
used in simulation
sensor/ temperature sensors: simulated, or a real
MAX31865 RTD amplifier over SPI
actor/ heater and stirrer drivers: simulated, a
Hendi induction hob (serial protocol), or a
Pololu 1376 stirrer motor controller
sud.py mash-schedule sequencer; starts out empty,
a client loads one of several sude/*.json
schedules onto it and it drives the
temperature controller from it
sud_forecast.py predicts how long a loaded schedule will
actually take by simulating it with the
same kind of plant/controller (and params)
as the real server - see "Forecast vs.
actual duration" below
tasks/ Async tasks (one per component) that poll
hardware/sim state at a fixed interval and
publish changes through the message dispatcher.
ws/ Minimal WebSocket pub/sub layer: server
(single- and multi-user), client, and a
keyed message dispatcher (e.g. "Sensor",
"Heater", "TempCtrl", "Stirrer", "Pot", "Sud").
tracer.py Logs traced variables to .mat files (for
offline analysis/tuning in MATLAB/Octave,
see results.m / results_tc.m).
scripts/demos/ Standalone, eyeballed-plot demos (matplotlib)
exercising components/* in isolation — pid/,
plant/, and a full mash-schedule run in sud/.
Not used at runtime; run directly, e.g.
`python -m scripts.demos.sud.demo_sud`.
sude/ Mash schedules ("Sud" = brew/wort), each a
JSON list of "ramp"/"hold" steps with target
temperature/heat rate or hold duration,
per-step stirrer timing, and optional pause
for user confirmation.
config.json.templ Configuration template (real hardware).
config.json.sim Configuration template for simulation mode.
```
### Data flow
The server (`server/brewpi.py`) loads `config.json`, builds the configured
sensor/heater/stirrer/controller via factories (`*Factory.create(name, ...)`)
based on the `Controller` section (`sensor_name`, `heater_name`,
`stirrer_name`, `pid_type`, or `"sim"` for any of them), and connects their
outputs to each other's inputs via `set_on_changed` callbacks, e.g.:
- sensor temperature → temperature controller's `theta_ist`
- temperature controller output `y` → heater power
- heater effective power → pot model power (simulation only)
Each component runs inside its own `ATask` at a configurable interval
(`Controller.dt`, scaled by `sim_warp_factor`) and pushes state changes onto a
keyed WebSocket channel that any connected client can subscribe to and send
commands back on (e.g. `{"TempCtrl": {"Soll": {"Temp": 65}}}`).
### State replay on connect
A fundamental assumption baked into this protocol: a client never knows what
it's connecting to. The GUI might be the first thing to ever talk to a
freshly started server, or it might connect to one that's been running
unattended for hours, mid-brew, with a schedule already loaded and paused
partway through step 4. A client can also disconnect (network hiccup, the
user closing and reopening the GUI) and reconnect later, at any server state,
with no special "resume" handshake. In every one of these cases, the client
must end up with the *same* full picture of current state as one that's been
connected the whole time - not just whatever changes happen after it joins.
This is what `WsServerMultiUser.global_state` (`ws/server/ws_server_multi_user.py`)
exists for: every message ever broadcast over the dispatcher is merged into
it (`ws/user.py`'s `update()`), keyed exactly like the live channels are.
Subscribing to a channel (`{"+": "Sud"}`, sent automatically for every
channel on connect - see `MessageDispatcherSync(auto_subscribe=True)`)
doesn't just start a future stream of changes; the server immediately
replays the *entire* currently-known state for that channel in one message,
via `User.send()`'s `dpath` search over `global_state`. A reconnecting
client gets exactly the same replay a brand-new one would, because the
server doesn't distinguish between the two cases at all - there's no
"first connect" special-casing anywhere in this path.
The corollary, easy to miss: this mechanism has no concept of "transient" -
anything sent through it is implicitly durable state that will be replayed
to whoever connects next, even long after the fact. A one-shot *event*
(something that happened, rather than something that's currently true) has
to be modeled as a value that gets explicitly cleared back to neutral right
after sending, or it will linger in `global_state` and get handed to some
unrelated future client as if it just happened to them. `tasks/sud.py`'s
`SudTask.recv()` hit this directly: a `Load` rejected because a run is in
progress sends `{"Sud": {"Error": "..."}}` to explain why, immediately
followed by `{"Sud": {"Error": None}}` to clear it - without that second
send, every client connecting afterwards, no matter how much later or how
unrelated to the rejected `Load`, would immediately see that same stale
error on connect. The client mirrors this by treating a falsy `Error` as a
no-op rather than an empty dialog.
## Requirements
- Python 3.8+
- Server (`server/requirements.txt`): `numpy`, `scipy`, `matplotlib` (only
for `scripts/demos/*`, not the running server itself), `websockets`,
`dpath`, and `pyserial`/`spidev` if using real hardware backends.
- Client (`client/requirements.txt`): `PyQt5`.
Install with:
```bash
pip install -r server/requirements.txt
pip install -r client/requirements.txt
```
## Running
1. Copy a config template to `config.json` next to `brewpi.py` and adjust it.
Use `config.json.sim` to run entirely in simulation (no hardware needed),
or `config.json.templ` as a starting point for real hardware (set the
heater/stirrer serial ports and sensor type).
2. Start the server:
```bash
cd server
./brewpi.py
```
This serves the WebSocket on `ws://0.0.0.0:8765`.
3. Start the GUI client and connect to the server's URI:
```bash
cd client
./brewpi_gui.py
```
`server/brewpi.sh` shows how this is wired up to run under a `virtualenvwrapper`
environment (`$WORKON_HOME`/`$BREWPI_HOME`) on a Raspberry Pi-style deployment.
## Mash schedules
Files under `sude/` describe a brew's mash schedule ("Sud"): `pot_mass`,
`pot_material` (fixed for the whole brew), and a `steps` list. Every step
ramps toward its `temperature`, then optionally holds:
- ramping toward `temperature` at `"ramp": {"rate": ...}` (°C/min) is not
opt-in - it happens for every step, for as long as the temperature
controller's own gap-tracking FSM (`components/pid/temp_controller_base.py`)
says the gap actually warrants it; a step omitting `temperature` simply
has no new target of its own and keeps whatever the previous step left
running, so the ramp resolves immediately. `ramp.rate` itself still comes
from `default.step.ramp` unless a step overrides it.
- a hold — `"hold": {"duration": ...}` — hold the current target for
`duration` minutes (omit/`0` for an immediate step) - this part stays
opt-in: only a step that specifies `hold` gets a hold-duration phase
after the ramp.
A step with both ramps to `temperature` and then holds there for `duration`
- useful for a mash rest ("ramp to 63°C, then hold 40 min") without needing
two separate schedule entries. `user_wait_for_continue`/`user_message`
(below) apply once the whole step is done, i.e. after the hold phase if
there is one.
Both `ramp` and `hold` carry a `stirrer` block (`speed`, `interval_time`,
`on_ratio`): `interval_time: 0` runs the stirrer continuously at `speed`;
`interval_time > 0` pulses it on a period of `interval_time` seconds, on for
`on_ratio * interval_time` of it. A step may also set
`user_wait_for_continue: true` (with an optional `user_message` to prompt
the user with) to pause for confirmation once the step completes, e.g. to
add malt or check gravity, instead of advancing immediately.
Each step also has its own `grain_mass`/`water_mass` (defaulted like
everything else from `default.step`), since both change over the course of
a brew — e.g. malt going in partway through, or water boiling off. Pass a
step's `grain_mass`/`water_mass` to `Sud.derive_plant_params()` (together
with the brew-wide `pot_mass`/`pot_material`) to get that step's lumped
`Pot` `M`/`C`; `scripts/demos/sud/demo_sud.py` recomputes and re-applies
these (via `Pot.set_thermal_params()`/`TempController.set_model_params()`,
on both the real plant and the controller's Smith-predictor model) every
time the current step changes, instead of deriving them once at startup.
The top-level `default.step` object gives every field above (including the
nested `ramp`/`hold`/`stirrer` blocks) a default value; a step in `steps`
only needs to specify the fields it overrides — anything it omits is filled
in from `default.step`, recursively. `ramp` is always defaulted in this way
(every step ramps); `hold` stays opt-in - only present if the step specifies
it. `temperature` is the one exception: it's never defaulted from
`default.step` (that's just inert template filler in every `sude/*.json`) -
a step either specifies its own, or has none and doesn't change the target.
The server always runs a `Sud` (`components/sud.py`), starting out empty (no
schedule) - `sude/` can hold several schedule files, and a client loads one
of them onto the running `Sud` (`{"Sud": {"Load": <sud.json contents>}}`),
kept purely in memory until replaced by another Load or the server
restarts; nothing is ever read from or written to `sude/*.json` by the
server itself, that's on the client (e.g. the GUI's File menu). `Sud`
resolves each raw step against `default.step` and tracks the current
(resolved) step, while `tasks/sud.py`'s `SudTask` drives the temperature
controller's `theta_soll`/`heatrate_soll` toward each step's target
(advancing once the controller's own FSM reports the gap closed -
`TempControllerBase.is_holding()` - rather than a separate ad hoc
tolerance), counts down hold steps' `duration`, and applies each step's
`stirrer` block (`interval_time`/
`on_ratio` map directly onto the stirrer's cycle time/duty cycle). It
exposes progress (current step, remaining hold time, elapsed run time -
tick-counted, not derived from wall-clock time times the warp factor, since
the latter is only nominal and drifts under real scheduling load - state,
any `user_message`) on the `"Sud"` WebSocket channel and accepts
`{"Sud": {"Start": true}}` to begin the schedule (also restarts one that's
already finished), `{"Sud": {"Pause": true}}` to freeze progress without
losing it, `{"Sud": {"Confirm": true}}` to acknowledge a pause and move to
the next step, and `{"Sud": {"Load": <sud.json contents>}}` to replace the
running schedule - refused (with an explicit, one-shot `Error` reply - see
"State replay on connect" above) while a run is already in progress, rather
than the client trying to guess that from its own view of the state.
The temperature controller has a master enabled switch (off by default -
see `components/pid/temp_controller_base.py`): `SudTask` enables it for as
long as a run is in progress (any state other than idle/finished) and
disables it again - forcing the heater output to 0 - the moment it stops or
finishes, handing control back to manual mode. The GUI's Manual tab has an
"Enabled" checkbox for driving this directly outside of a Sud run; while
disabled, its temperature/heat-rate setpoint controls are inactive.
### Forecast vs. actual duration
The GUI's Automatic tab shows up to three time estimates for a schedule:
- Immediately on load, a quick *naive* preview: walk the schedule assuming
every ramp instantly achieves and holds its declared `rate`
(`abs(delta)/rate`) and every hold lasts exactly its declared `duration`.
Deliberately approximate - it's just a placeholder until the next one
arrives a moment later.
- A *simulated* estimate, computed server-side (`components/sud_forecast.py`'s
`SudForecastEstimator`) by actually running the schedule through a
throwaway `Sud`/`Pot`/temperature-controller trio built from the same
params (`pid_type`, `TempCtrl` gains, plant params, ambient, heater max
power) the real server uses - a multi-hour brew simulates in well under a
second since it's pure CPU-bound iteration, no real time/IO involved. This
replaces the naive preview a moment after a schedule is loaded.
- Once running, a *dynamic* one: the already-elapsed part is the actual
measured trace, plotted against `Sud.elapsed` (see "State replay on
connect" above) rather than wall-clock time, and only the remaining steps
are re-projected from the live temperature/step/`hold_remaining` each
tick.
The naive estimate is structurally optimistic - it assumes the plant
instantly tracks the declared rate and that "reached" is instant once the
math says so, when the real PID cascade (`theta_err -> pid_hold ->
heatrate_soll -> pid_heat -> heater power -> actual heat rate`) needs real
time to spin up and settle into the controller's own `HOLD` state
(`TempControllerBase.is_holding()`, gated by the `HoldHeat`/`HoldCool`
thresholds). The simulated estimate accounts for that (it's driven by the
real controller dynamics), so it's normally
within a few percent of how the dynamic one settles once a run actually
finishes - if the two diverge by far more than that, suspect a real control
issue rather than a forecasting one (see the `HoldCool` threshold note in
`components/pid/temp_controller_base.py` - too tight a threshold there once
caused exactly this kind of large, otherwise-unexplained gap, by making the
controller chatter in and out of `COOL` and never actually converge). Rule
out a *display* artifact first, though: the dynamic trace used to position
itself using wall-clock time times the configured `sim_warp_factor`, but
that factor is only nominal - real asyncio/Python scheduling overhead means
the actually-achieved speedup runs measurably below it (confirmed: ~80x
instead of a configured 100x), which alone produced a divergence large
enough to look like a control problem. `Sud.elapsed` (tick-counted, immune
to wall-clock pacing) replaced that reconstruction for exactly this reason.
## Logging & analysis
`tracer.py` (and `tasks/tracer.py`) periodically dump traced signals
(temperatures, heat rates, heater power) to `logs/*.mat` files, which can be
loaded in MATLAB/Octave — see `results.m` and `results_tc.m` — to evaluate
and tune the controller.
`server/brewpi.py` also mirrors everything it prints (every component's
`print()`-based status/debug output) to a plain text log,
`logs/brewpi.<timestamp>.log`, alongside the `.mat` traces - useful for
post-mortems without needing to have been watching the console live.