Files
brewpi/docs/overshoot_hold_windup.md
T
jensandClaude Sonnet 5 2d032bf9b2 fix: allow gentle negative HOLD floor to fix steady-state overshoot
Testing the windup-fix rework live against sude/sud_0030.json surfaced a
distinct bug: a HOLD overshoot from grain-fill-in cooling never decayed -
process_pid()'s pid_outer_y floor clamped to exactly 0.0, so pid_inner
fought the pot's own ambient loss to hold the overshot temperature flat
instead of declining back to setpoint (see docs/overshoot2.png).

Replace the hardcoded 0.0 floor with a configurable Outer.y_hold_min
(default 0.0, backward compatible), set to -0.1 in config.json, both
.tpl templates, and the demo scripts - small enough to avoid
reintroducing the bb5af3c limit cycle while letting HOLD request a
gentle decline matching passive ambient cooling.

Adds TestHoldOvershootRecoversToSetpoint and documents the finding in
docs/overshoot_hold_windup.md's Follow-up section and
components/pid/TODO.md. Confirmed against a live sud_0030 re-run, not
just the unit test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M2ierBoxW3v7nUbDw3M2pE
2026-07-06 08:19:10 +02:00

17 KiB

HOLD-state overshoot from a disturbance: root cause and fix

Investigation of the overshoot visible in docs/overshoot.png: a cold-water disturbance during a steady HOLD at 30°C caused theta_ist to overshoot to ~30.5°C and stay there, rather than settling back at soll. This document records the diagnosis, the fix that shipped for it, and the test coverage added alongside it. Numbers in the incident summary below are pulled from logs/log_latest.json; names in that section (pid_hold/pid_heat) are the pre-fix names — see "Rename for honesty" below for what they became.

Incident summary

  • Steady HOLD at 30.0°C. Cold water added → theta_ist (Smith-corrected) dipped to ~29.49°C (diff ≈ 0.51°C).
  • diff never reached Thresholds.HoldHeat (1.0°C, temp_controller_fsm.py), so the FSM never left HOLD — the entire episode happened inside what the system still considered "holding".
  • The outer loop (pid_hold, Hold: {kp:0.6, ki:0, kd:0, kt:0} — pure proportional) tracked 0.6 * max(0, diff) exactly, tick for tick, peaking at y ≈ 0.31 — nowhere near its y_max=1.0 ceiling. Confirmed directly against logged samples: rate_soll == 0.6*diff to the last digit throughout.
  • The inner loop (pid_heat, Heat: {kp:0.08, ki:0.02, kd:0, kt:1.5}) then chased that heatrate_soll target, driving commanded power from a baseline ~500W burst-cycle up to ~2600W (y ≈ 0.74 of the 3500W max, see tasks/heater.py:58-59: power = get_power_max() * y).
  • Plant dead time (Td=17s) and thermal mass (M=27.96 kg, C≈3403 J/(kg·K), from PlantParams in the log) meant heat from that burst kept arriving after theta_ist had already crossed back over soll. By the time rate_soll had fallen back to exactly 0 (outer loop asking for zero heat rate), power_set was still 1442 W, decaying only gradually over the next ~35s (1193 → 979 → 817 → 632 → 487 → 337 → 142 W) rather than snapping to 0.
  • No active cooling actuator exists; the passive loss coefficient (L=0.2 W/(kg·K)) gives a time constant C/L ≈ 17,000s (~4.7h), so the resulting ~0.4°C plateau above soll does not visibly decay within any reasonable observation window.

Root cause: the inner loop's integral term (yi) accumulated during the ~130s the outer loop demanded a real (if modest) heat rate, and nothing bounded that accumulation — the existing anti-windup in Pid.process() (components/pid/pid.py:45-48) is back-calculation that only corrects yi when the combined output y is actually clamped at y_min/y_max. Since y peaked at ~0.74, well under the 1.0 ceiling, awu was 0 the entire time and the anti-windup mechanism never engaged. The integrator just had to unwind naturally against real (delayed) negative error, which took ~35+ seconds after the outer loop had already zeroed its target.

Rejected approach: a flat yi_max clamp on the inner loop

The first fix considered was bounding yi directly, with one constant for the inner loop regardless of state. Numeric check against the real plant model kills this as a standalone fix:

scenario P needed y needed
1.5 K/min ramp @ 30°C 2435 W 0.696
1.5 K/min ramp @ 50°C 2547 W 0.728
1.5 K/min ramp @ 66°C 2636 W 0.753
the actual HOLD disturbance (peak) 2600 W 0.74
steady HOLD loss compensation @ 66°C ~257 W ~0.073

At steady state (err=0), kp*err contributes nothing — the entire y≈0.7-0.75 needed to sustain a genuine 1.5 K/min ramp has to come from yi alone, for as long as the ramp lasts. That's the same range the disturbance transient itself peaked at, but far above what real steady-HOLD loss compensation ever needs (~0.07-0.15 at realistic brew temperatures). A single constant can't serve both: tight enough to matter for the disturbance (well under ~0.74) permanently starves a real ramp of the ~0.7 it needs; loose enough not to interfere with ramps (~0.75+) never engages during the disturbance at all. Rejected as a single global constant.

An FSM-gating alternative (freeze the inner loop's output to 0 while in HOLD unless explicitly engaged by command or a new threshold) was also sketched, but left an open question about how a new engage threshold should relate to the existing HoldHeat FSM threshold, and needed extra engage/disengage hysteresis to avoid chatter. Superseded by the plan below, which reaches the same effect with less new machinery.

Refined plan: split the clamp by which state is driving the same loop,

not by freezing the loop

The insight above — real HOLD-time demand (~0.07-0.15) and real ramp demand (~0.7-0.75) occupy clearly different ranges — means a per-state parameter set on the same PID instance solves this more cleanly than gating the loop on/off: same accumulated yi/d state carried across HOLD↔HEAT transitions (bumpless transfer for free, no explicit reset, no engage/ disengage hysteresis to tune), just a different yi_max ceiling depending on which state is currently active.

Rename for honesty

The pre-fix names didn't match what actually ran when:

  • pid_hold already ran unconditionally every tick regardless of state (process_pid()'s first line, temp_controller_base.py:137) — it's the outer loop, not "the HOLD-state PID". Renamed to pid_outer.
  • pid_heat already ran in both HOLD and HEAT (only IDLE/COOL skip it, temp_controller_base.py:154-163) — it's the inner loop for the heating direction. Renamed to pid_inner.
  • pid_cool only ever ran in COOL — no HOLD-time ambiguity, since a disturbance that pushes temp above soll during HOLD is still handled by pid_inner (the outer loop's max(0.0, pid_outer_y) floor sends heatrate_soll to 0, and pid_inner reacts to the resulting negative heatrate_err — the state only escalates to real COOL past Thresholds.HoldCool). Renamed to pid_inner_cool for symmetry, kept as its own instance — the explicit reset() calls when crossing between heat-direction and cool-direction states (temp_controller_fsm.py:104,108,115,123) stay exactly as they were; there is no reason to share integrator state across a heater/chiller boundary.

Config: Hold/Heat/CoolOuter / Inner.{Heat,Hold,Cool}

"TempCtrl": {
  "pid_type": "Smith",
  "beta": 0.9,
  "Outer": { "kp": 0.6, "ki": 0.0, "kd": 0.0, "kt": 0.0 },
  "Inner": {
    "Heat": { "kp": 0.08, "ki": 0.02, "kd": 0.0, "kt": 1.5 },
    "Hold": { "kp": 0.08, "ki": 0.02, "kd": 0.0, "kt": 1.5, "yi_max": 0.3 },
    "Cool": { "kp": 0.08, "ki": 0.02, "kd": 0.0, "kt": 1.5 }
  },
  "Thresholds": { "...": "unchanged" }
}
  • Inner.Heat keeps today's Heat gains, no yi_max (or a very loose one) — a real ramp must be able to reach y≈0.75.
  • Inner.Hold starts as a copy of the same gains, with yi_max≈0.2-0.3 added — comfortably above realistic steady-loss compensation (~0.07-0.15) but well below what turned a 0.5°C dip into a 2600W burst.
  • Inner.Cool is Cool's existing gains, moved under Inner purely for structural consistency — introduced now, not because we've observed a cooling-side incident. pid_inner_cool never runs during HOLD, so it doesn't need its own Hold variant the way Heat does; one params block is enough.
  • This was a breaking config change — no backward-compat shim for the old flat Hold/Heat/Cool keys (per the "no compat hacks" convention). Every deployed config.json needed migrating, not just the repo's config-real.json.tpl/config-sim.json.tpl templates.

Code changes (implemented)

  1. components/pid/pid.py — the symmetric yi_max clamp lives inside process() (line 40): self.yi = max(-yi_max, min(yi_max, self.yi)) when self.params.get('yi_max') is set, applied right after accumulating yi and before it's summed into y.
  2. components/pid/temp_controller_fsm.pyself.pid_holdself.pid_outer, self.pid_heatself.pid_inner, self.pid_coolself.pid_inner_cool (constructor at lines 33/34/40, all reset() call sites and comments at lines 11-12, 83, 87-89, 99, 104, 108, 111, 115, 119, 123).
  3. components/pid/temp_controller_base.py:
    • set_params() (lines 30-37): self.pid_outer.set_params(params['Outer']); stores self._inner_heat_params = params['Inner']['Heat'] and self._inner_hold_params = params['Inner']['Hold'] for the per-tick lookup below; self.pid_inner_cool.set_params(params['Inner']['Cool']).
    • process_pid() (lines 136-163): pid_hold_ypid_outer_y, and the self.pid_hold.process(...) call. In the combined HOLD/HEAT branch (lines 159-163), the active param set is selected before processing:
      else:
          inner_params = self._inner_heat_params if self.state == States.HEAT else self._inner_hold_params
          self.pid_inner.set_params(inner_params)
          self.pid_inner.process(heatrate_err, -self.heatrate_ist)
          self.y = self.pid_inner.get_y()
      
      set_params() is a cheap dict-reference assignment (pid.py:22-23), so calling it every tick has no meaningful cost. Because kp/ki/kd/ kt are identical between Inner.Heat and Inner.Hold in the shipped config, switching the active set at a HOLD↔HEAT transition changes no term of y at that instant — only the yi_max ceiling going forward, with one caveat noted in Status below.
  4. Config filesconfig.json, config-real.json.tpl, config-sim.json.tpl restructured into Outer/Inner.{Heat,Hold,Cool} as above. The inline "Cool": {...} dicts in scripts/demos/pid/ demo_temp_controller_smith.py, demo_temp_controller.py, and scripts/demos/sud/demo_sud.py got the same restructure.
  5. utils/replay_sim.py_apply_gain_overrides() and the CLI flag loop now iterate the shared GAIN_SECTIONS = (('Outer','outer'), ('Inner.Heat','inner-heat'), ('Inner.Hold','inner-hold'), ('Inner.Cool','inner-cool')), with nested dict access for the Inner.* entries. The params print loop and _infer_heatrate_soll_set()'s docstring were updated to match (pid_outer.get_y() instead of "the hold PID").
  6. components/pid/TODO.md — the windup entry is marked [x] and points at this section.

Tests

Stdlib unittest, no pytest — tests/components/pid/test_pid.py and test_temp_controller_closed_loop.py, discoverable via python -m unittest discover -t . -s tests/components/pid (or -s tests for the whole repo suite). Simpler than the FSM-gating plan's test plan would have needed: no engage/disengage hysteresis or threshold-relationship behavior to cover, since the loop is never turned off — only its yi_max ceiling changes with state.

A. Unit-level, isolated Pid (test_pid.py) — no plant involved. Feeds a synthetic err sequence shaped like the incident (positive heatrate_err = 0.3 held for 130 ticks, matching the outer loop's real demand during the disturbance, then a flat -0.3 tail for 200 ticks, matching the real rate_soll - rate_ist gap once the outer loop had zeroed its target) into two Pid instances with identical gains, one with yi_max set (the Inner.Hold case) and one without (Inner.Heat). Asserts: recovery time (ticks after the error goes negative until y drops back under a small threshold) is measurably shorter when clamped; yi never exceeds the configured bound; the unclamped instance's yi does exceed it (sanity-checks the test itself isn't vacuous).

B. Closed-loop (test_temp_controller_closed_loop.py) — real Pot(dt) plant with M=27.96, C=3403.43, L=0.2, Td=17, ambient 20°C, driven by a real TempController(Smith) with the shipped Outer/Inner.* gains. Controller's own y feeds back into the plant each tick (unlike utils/replay_sim.py's open-loop observe-only mode):

  • Disturbance case: hold at 30°C until settled, knock plant.temp down 0.5°C to emulate the cold-water event, keep ticking for 600 more ticks — confirms the state stays in HOLD throughout, matching the incident. Asserts peak overshoot above 30.0°C is measurably smaller with Inner.Hold's yi_max set than with an unclamped copy of Inner.Heat.
  • Ramp case: commands a genuine 1.5 K/min ramp to 40°C (state reaches HEAT) and asserts the sustained heat rate actually exceeds 1.4 K/min — guards against reintroducing the flat-clamp regression from the rejected approach above.
  • Transition case: drives a HOLD→HEAT→HOLD sequence and asserts the y step at either transition stays under 0.1 — see the retroactive-clamp caveat in Status below for why this isn't a stricter "no discontinuity" assertion.

Status

Implemented: pid.py's yi_max clamp, the pid_outer/pid_inner/ pid_inner_cool rename, the Outer/Inner.{Heat,Hold,Cool} config restructure (config.json, both .tpl templates, and the three demo scripts), and utils/replay_sim.py's matching CLI-flag/print-loop rename. Tests added under tests/components/pid/ (test_pid.py for the isolated Pid clamp behavior, test_temp_controller_closed_loop.py for the closed-loop disturbance/ramp/transition cases) — all passing.

One subtlety found while writing the closed-loop transition test that this plan didn't anticipate: Inner.Hold's yi_max clamp applies retroactively. If a sustained HEAT ramp pushes yi above the Hold ceiling before the HeatHold threshold fires, the very next tick after the HEAT→HOLD transition clamps yi back down immediately, producing a small (~0.07 in testing, well below the pre-fix disturbance's ~0.74 peak) step in y rather than the fully bumpless transfer described above. Not addressed here — flagged for awareness, not a blocker.

Follow-up: steady-state overshoot in HOLD (Outer.y_hold_min)

Found while testing the rework against a real Sud run (sude/sud_0030.json, logs/log_20260706T074658_Sud-0030.json). Visible in docs/overshoot2.png (theta_ist vs theta_soll across the run): at both the 55°C and 63°C rests, theta_ist overshoots the step and then plateaus above theta_soll for the rest of the hold instead of converging back down — most clearly at the first rest, where it settles at ~55.5-55.6°C against a 55.0°C target.

A grain-fill-in disturbance during "1. Rast" (mash-in rest, HOLD at 55°C) pushed temp_ist to a ~0.4°C overshoot — well under HoldCool's 1.0 threshold, so the FSM stayed in HOLD throughout, same as the transient windup case above. But this overshoot never decayed: temp_ist sat in a 55.30-55.44 band for the rest of the 20-minute hold instead of converging back to 55.0. Distinct failure mode from the transient windup fixed above (that one unwound over ~35s; this one was flat/permanent for as long as the hold lasted).

Root cause: process_pid()'s HOLD-state floor on pid_outer_y (added in bb5af3c to break a limit cycle - see the History section) clamped to exactly 0.0. Once temp_ist > temp_soll, that floor forces heatrate_soll = 0, i.e. "hold flat" - and pid_inner then actively fights the pot's own ambient heat loss to keep the overshot temperature flat, rather than being allowed to request a genuine decline back toward setpoint. With Outer.ki = 0, there's no integral action to null the resulting steady-state error any other way, so the offset persists for the rest of the hold.

Fix: replaced the hardcoded 0.0 floor with a configurable Outer.y_hold_min (default 0.0, so configs that don't set it keep the old behavior), set to -0.1 in config.json, both .tpl templates, and the three demo scripts. A small negative floor lets HOLD ask for a gentle decline that roughly matches passive ambient cooling, rather than the fully unclamped [-1, 1] range that caused the original limit cycle (a large negative heatrate_soll asks for a decline steeper than passive loss can deliver, pinning power at 0 for an extended stretch and producing a hard undershoot/rebound). Test coverage added: TestHoldOvershootRecoversToSetpoint in tests/components/pid/test_temp_controller_closed_loop.py reproduces the sud_0030 disturbance, asserts the old flat-clamp behavior still fails to recover (regression guard) and the new floor converges close to setpoint, plus a guard that the recovery doesn't undershoot by more than the injected disturbance itself (i.e. doesn't reintroduce the bb5af3c limit cycle).

-0.1 was chosen from the real Sud's plant params (Pot.mass=5.96 + water_mass=22 ≈ the test harness's M=27.96, L=0.2): passive ambient loss at a ~35°C delta works out to roughly 0.1-0.15 K/min, so heatrate_soll_set * -0.1 lands in that same ballpark for a typical heatrate_soll_set of ~1.0-1.5 K/min. Not derived from first-principles tuning - may need adjustment per installation, same as the other PID gains.

Confirmed against a live re-run of sud_0030 with the fix applied, not just the unit test above.

Architecture diagram

A full signal-flow diagram of the cascade (Outer/Inner.* PIDs, FSM state gating, Smith-predictor feedback) exists as a Claude Artifact: https://claude.ai/code/artifact/a32e4752-b4a6-4153-b344-eb2423eb6512 — this is a session-scoped link, not a durable one, so it may not resolve for everyone with repo access. docs/fsm_states.png is a static screenshot of just the diagram's FSM-states panel, checked in as the durable copy.