Records the root cause of the cold-water overshoot in docs/overshoot.png (integral windup in pid_heat with no anti-windup engagement, since y never saturated) and the refined fix plan: gate pid_heat by FSM state instead of a flat yi_max clamp, which was rejected after the numbers showed it would also cripple legitimate high-rate ramps. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KefnwbGDM8CGrhb4sFhVq9
8.0 KiB
HOLD-state overshoot from a disturbance: root cause and fix plan
Investigation of the overshoot visible in docs/overshoot.png: a cold-water
disturbance during a steady HOLD at 30°C caused theta_ist to overshoot to
~30.5°C and stay there, rather than settling back at soll. This document
records the diagnosis and the (not yet implemented) fix plan. Numbers below
are pulled from logs/log_latest.json.
Incident summary
- Steady HOLD at 30.0°C. Cold water added →
theta_ist(Smith-corrected) dipped to ~29.49°C (diff ≈ 0.51°C). diffnever reachedThresholds.HoldHeat(1.0°C,temp_controller_fsm.py), so the FSM never leftHOLD— the entire episode happened inside what the system still considered "holding".pid_hold(outer loop,Hold: {kp:0.6, ki:0, kd:0, kt:0}— pure proportional) tracked0.6 * max(0, diff)exactly, tick for tick, peaking atpid_hold_y ≈ 0.31— nowhere near itsy_max=1.0ceiling. Confirmed directly against logged samples:rate_soll==0.6*diffto the last digit throughout.pid_heat(inner loop,Heat: {kp:0.08, ki:0.02, kd:0, kt:1.5}) then chased thatheatrate_solltarget, driving commanded power from a baseline ~500W burst-cycle up to ~2600W (y ≈ 0.74of the 3500W max, seetasks/heater.py:58-59:power = get_power_max() * y).- Plant dead time (
Td=17s) and thermal mass (M=27.96 kg,C≈3403 J/(kg·K), fromPlantParamsin the log) meant heat from that burst kept arriving aftertheta_isthad already crossed back over soll. By the timerate_sollhad fallen back to exactly0(outer loop asking for zero heat rate),power_setwas still 1442 W, decaying only gradually over the next ~35s (1193 → 979 → 817 → 632 → 487 → 337 → 142 W) rather than snapping to 0. - No active cooling actuator exists; the passive loss coefficient
(
L=0.2 W/(kg·K)) gives a time constantC/L ≈ 17,000s (~4.7h), so the resulting ~0.4°C plateau above soll does not visibly decay within any reasonable observation window.
Root cause: pid_heat's integral term (yi) accumulated during the
~130s the outer loop demanded a real (if modest) heat rate, and nothing
bounded that accumulation — the existing anti-windup in Pid.process()
(components/pid/pid.py:45-48) is back-calculation that only corrects yi
when the combined output y is actually clamped at y_min/y_max. Since
y peaked at ~0.74, well under the 1.0 ceiling, awu was 0 the entire
time and the anti-windup mechanism never engaged. The integrator just had to
unwind naturally against real (delayed) negative error, which took ~35+
seconds after the outer loop had already zeroed its target.
Rejected approach: flat yi_max clamp on pid_heat
The first fix considered was bounding yi directly
(self.yi = max(-yi_max, min(yi_max, self.yi)) in Pid.process(),
independent of the existing y-clamp anti-windup). Numeric check against
the real plant model kills this as a standalone fix:
| scenario | P needed | y needed |
|---|---|---|
| 1.5 K/min ramp @ 30°C | 2435 W | 0.696 |
| 1.5 K/min ramp @ 50°C | 2547 W | 0.728 |
| 1.5 K/min ramp @ 66°C | 2636 W | 0.753 |
| the actual HOLD disturbance (peak) | 2600 W | 0.74 |
At steady state (err=0), kp*err contributes nothing — the entire
y≈0.7-0.75 needed to sustain a genuine 1.5 K/min ramp has to come from
yi alone, for as long as the ramp lasts. That's the same range the
disturbance transient itself peaked at. So:
yi_maxset low enough to meaningfully shorten the observed hangover (well under ~0.74, e.g. ~0.3) would permanently cap a real 1.5 K/min ramp far below its commanded rate — not a transient dip, a persistent, uncorrectable shortfall for the entire ramp.yi_maxset high enough not to interfere with legitimate ramps (~0.75+) sits at or above what the incident already peaked at, so it never engages and does nothing for the hangover.
A flat magnitude clamp cannot distinguish "leftover integral from an already-resolved disturbance" from "integral correctly holding up a real ongoing ramp" — in this plant they occupy the same output range. Rejected.
Refined plan: gate pid_heat by FSM state instead
Bound (or freeze) the rate at which pid_heat reacts based on FSM state,
not a constant on yi, since real sustained-rate demand only ever happens
in HEAT; in HOLD any large yi is by definition windup, since the
target rate there is always ~0 outside of a real disturbance.
- New flag
heat_loop_active, scoped toHOLDonly, inTempControllerFsm.is_holding()staysstate == HOLDregardless — Sud step "have we reached target" logic is unaffected. - While
heat_loop_activeis False inHOLD: forcey=0, skippid_heat.process()entirely (freeze it, don't reset — same bumpless-resume convention already used forHOLD→HEAT, seetemp_controller_fsm.py:99-101).pid_holdkeeps running every tick regardless (it's pure-P, no windup risk, and its output is needed to evaluate the engage condition below). - Engagement (
False→True), checked each tick while inHOLD:- Explicit command — a new
engage_heat_loop()/force_heat()call, wired to a manual UI action and/or auto-fired fromset_theta_soll()when the new target is a genuine change (not the same value re-sent every tick). diff >= Thresholds.HoldHeatEngage(new config key).
- Explicit command — a new
- Disengagement: once
diffsettles back inside the existingHoldHeat/HoldCoolband, held for a short dwell time (a few seconds), to avoid chatter right at the boundary. - Config/back-compat: new threshold merges into
DEFAULT_THRESHOLDSthe same way existing ones do (temp_controller_base.py:30).
Open design question (unresolved)
HoldHeatEngage needs to sit below the existing HoldHeat/HeatHold
FSM threshold (currently 1.0°C) to ever fire from inside HOLD — if
HoldHeat stays at 1.0, the FSM fully escalates to HEAT (where
pid_heat always runs anyway) before any larger engage threshold is ever
reached. Two ways to resolve, not yet decided:
- Raise
HoldHeat/HeatHoldtoo, widening the HOLD↔HEAT band soHOLDcovers the full deadband and the new flag is the only gate within it. - Keep
HoldHeatat1.0, set the engage threshold below it (e.g. ~0.7-0.8°C) — loop only goes offline for small disturbances, engaging before the FSM would escalate to fullHEATanyway.
Test plan (not yet implemented)
Stdlib unittest, no pytest — tests/components/pid/test_pid.py,
discoverable via python -m unittest discover -t . -s tests/components/pid.
A. Unit-level, isolated Pid — no plant involved. Feed a synthetic
err sequence shaped like the incident (positive heatrate_err ≈ 0.3 held
for ~130 ticks, then decaying/negative tail, matching the real
rate_soll - rate_ist pulled from the log) into pid_heat with the loop
gated (frozen while "HOLD"/inactive) vs ungated (today's behavior). Assert:
recovery time (ticks after error goes negative until y drops back under a
small threshold) is measurably shorter when gated; yi never grows during
the frozen window.
B. Closed-loop, self-contained synthetic scenario — no dependency on
the multi-MB log file. Real numbers from PlantParams/config.json:
Pot(dt) with M=27.96, C=3403.43, L=0.2, Td=17, ambient from
config.json; TempController via PidFactory.create('Smith', dt) with
the real Hold/Heat/Cool gains. Run closed-loop (controller's own y
drives the plant, unlike utils/replay_sim.py's open-loop observe-only
mode): hold at 30°C until settled, knock plant.temp down ~0.5°C to
emulate the cold-water event, keep ticking for several minutes. Run twice —
gating on vs off — and assert peak overshoot above 30.0°C is measurably
smaller with gating, while a separate run driving a genuine 1.5 K/min ramp
confirms gating does not reduce the sustained ramp rate (guards against
reintroducing the flat-clamp regression above).
Status
Plan only — nothing in this document has been implemented yet.