Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
14 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ There is no source-code lint target or unit-test suite; ESPHome config validatio
- `keypad_protocol_types.md` — packet type catalogue (known vs inferred meanings)
- `arm_disarm_state_machine.md` / `output_select_state_machine.md` — state machine design rationale
- `protocol_investigations.md` — raw observations and open questions
- `known_quirks.md` — user-facing translation of the above: symptoms a user could actually notice (log warnings, delayed disarms), why they happen, and whether action is needed. Update it when an investigation finding has a user-visible symptom, not just internal byte-level detail.

Distinguish observed facts from inferred hypotheses when adding to these docs or to log messages; unknown packet meanings are labeled conservatively. `skills/protocol-reverse-engineering/skill.md` defines the analysis methodology used for trace work.

Expand Down
4 changes: 4 additions & 0 deletions components/crow_alarm_panel/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -248,3 +248,7 @@ logger:
logs:
crow_alarm_panel: DEBUG
```

Seeing an occasional `WARN` line or a disarm/arm that takes a few extra seconds? Check
[`docs/known_quirks.md`](docs/known_quirks.md) before assuming it's a new bug — most
recurring log noise from this component is already understood and benign.
9 changes: 9 additions & 0 deletions components/crow_alarm_panel/crow_alarm_panel.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -917,6 +917,15 @@ void CrowAlarmPanel::loop() {
}
break;
}
case RF_REMOTE_EVENT: {
if (data.size() < 5) {
ESP_LOGW(TAG, "RF remote event too short, discarding");
break;
}
ESP_LOGD(TAG, "[RF Remote %02x:%02x:%02x] Button event, code %02x:%02x [%02x.%s]", data[0], data[1],
data[2], data[3], data[4], type, format_hex_pretty(data).c_str());
break;
}
default:
ESP_LOGD(TAG, "Unknown [%02x.%s]", type, format_hex_pretty(data).c_str());
break;
Expand Down
5 changes: 5 additions & 0 deletions components/crow_alarm_panel/crow_alarm_panel.h
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,11 @@ static const uint8_t OUTPUT_STATE = 0x50;
static const uint8_t CURRENT_TIME = 0x54;
static const uint8_t KEYPAD_PING = 0x23; // Observed recurring keypad keep-alive/ping traffic
static const uint8_t KEYPAD_REGISTRATION = 0xA0; // Keypad announce: [a0.address.00]; sent on power-up/reset

// RF remote button-press event, not tied to any keypad address. data[0..2] is a fixed per-remote
// identity; data[3..4] is a per-remote, per-button code (not portable across remote units — see
// docs/protocol_wire_format.md's 0x7C entry).
static const uint8_t RF_REMOTE_EVENT = 0x7C;
static const uint8_t BOUNDARY = 0x7E;
// static const uint8_t KEYPRESS = 0xD1; // This is from upstream, but doesn't get sent by arrowhead panels that I can see
static const uint8_t KEYPRESS = 0xA1;
Expand Down
51 changes: 51 additions & 0 deletions components/crow_alarm_panel/docs/arm_disarm_state_machine.md
Original file line number Diff line number Diff line change
Expand Up @@ -580,6 +580,56 @@ settle it.

**Practical takeaway:** no code change proposed. Still waiting on a disarm that coincides with a confirmed-stuck `CURRENT_TIME` state to further test the correlation — none occurred in this window.

### Update (2026-09-07): an 8h36m-armed disarm retried (2/6) with `CURRENT_TIME` confirmed *normal* the whole time — a genuine counter-example on the other side of the correlation

**Source:** the frigate long-term logger, window 2026-09-05 20:47 → 2026-09-07 05:14 UTC (~32.5h, spans an OTA reboot at 2026-09-05 22:22:48 UTC — docs/comment-only, no functional change, see `protocol_investigations.md`'s 2026-09-07 update).

**Observed facts — three arm/disarm cycles this window:**

| Disarm (UTC) | Initiator | Armed since | Duration armed | Retries | `CURRENT_TIME` at disarm |
| --- | --- | --- | --- | --- | --- |
| 09-06 01:17:02 | physical (`AAP Keypad`) | 09-05 23:26:39 (physical arm) | ~1h50m | none, instant | normal |
| 09-07 03:20:50 | integration (`disarm()`, `ESPHome Keypad`) | 09-06 18:44:23 (physical arm) | ~8h36m | **2/6, ~6.7s (backoff 1s→2s)** | normal, confirmed (surrounding `Controller time update`/glitch-recovery broadcasts all self-correct cleanly, no stuck-state broadcast anywhere near the retry) |
| 09-07 04:08:02 | integration (`disarm()`, `ESPHome Keypad`) | 09-07 03:21:37 (physical arm) | ~46m25s | none, instant | normal |

(A third arm event, 09-06 18:44:23, and its corresponding disarm above bracket the retried cycle; all three arms this window were physical-keypad-initiated, consistent with prior windows where most arms are physical and most disarms are the integration.)

**Inference (medium confidence — revises the 2026-09-03/09-05 correlation):** the middle row is a clean counter-example on the side of the correlation that had, until now, no exceptions: every previously-confirmed-normal-time disarm resolved instantly (2026-09-03 table: 2/2; 2026-09-05 update: 2/2), while every retry case had either confirmed-stuck time or unknown status. This is the first retry observed with `CURRENT_TIME` *confirmed normal* throughout. Combined with the existing 2026-08-31 counter-example on the other side (confirmed-stuck time, instant 3s-later disarm), the `CURRENT_TIME`-stuck correlation no longer looks like even a partial predictor — stuck time is neither necessary nor sufficient for a retry, across the full sample gathered so far. Duration-armed also doesn't split cleanly: this window's longest-armed case (8h36m) retried while a prior window's longest (5h26m) didn't, and this window's two instant cases (1h50m, 46m) bracket the retried one without an obvious threshold.

**Assumption (unverified):** what actually distinguishes the small number of retry cases from the larger number of instant ones is now back to being an open question rather than one with a leading candidate. No other signal (bus corruption rate, watchdog activity, time-of-day) has been checked yet for correlation with this specific retry.

**Practical takeaway:** no code change proposed — the retry/backoff machinery again resolved correctly (2/6, ~6.7s) regardless of cause. Given `CURRENT_TIME`-stuck no longer looks predictive, future sessions should stop treating it as the leading hypothesis and instead capture whatever else is observable at the moment of a retry (recent `ff.`/`fe.` corruption, watchdog trips, or anything else nearby in the log) to look for a different pattern from scratch.

### Update (2026-09-09): first data point for the "other signals" search — a retry ~5 minutes after a corruption-burst/watchdog-trip episode, vs. three bus-clean instant cycles

**Source:** the frigate long-term logger, window 2026-09-08 08:40 → 2026-09-09 05:09 UTC. Three arm/disarm sequences this window; the first two were confirmed by the user to be triggered by the official RF remote (distinct chime pattern: once on arm, twice on disarm) rather than a keypad or this integration — at the time of this entry, an `ARMED_STATE` broadcast considered in isolation couldn't distinguish remote-fob input from any other non-integration source, since both simply appear as Controller `ARMED_STATE` broadcasts with no accompanying keypress/command traffic. (A direct RF-remote signature, the `0x7C` frame preceding the broadcast, was decoded later — see `protocol_wire_format.md`'s `0x7C` entry — but wasn't yet identified when this window was analyzed.)

**Observed facts:**

| Sequence | Arm (UTC) | Disarm (UTC) | Duration armed | Retries | Nearby `ff.`/`fe.`/watchdog activity |
| --- | --- | --- | --- | --- | --- |
| 1 (remote) | 23:23:41 → confirmed 23:24:10 (09-08) | 23:25:00 | ~50s | none, instant | none within ±10min |
| 2 (remote) | 02:11:49 → confirmed 02:12:17 (09-09) | 02:12:30 | ~13s | none, instant | none within ±10min |
| 3 (integration) | `arm_away()` 03:48:43 → confirmed 03:49:11 | `disarm()` 04:43:55 | ~54m44s | **2/6, ~6.2s (backoff 1s→2s)** | clustered burst `04:37:40–42` (8 lines) → watchdog trip `04:38:40` → 3 repeated `Armed Away` re-broadcasts, all ending **~5min before** the `04:43:55` disarm call; nothing closer, nothing during the retry window itself |

**Inference (low confidence — one data point):** this is the first retry-associated case where a bus-corruption/watchdog episode occurred anywhere near the event, in either direction — the two remote-triggered cycles (bus-clean, instant) and the earlier `CURRENT_TIME`-normal instant cases from prior windows give no reason to expect corruption alone to force a retry, but this is also the first time anyone has actually looked. The ~5-minute gap is a meaningful caveat against reading too much into it: nothing in the doc's existing model gives an episode that already resolved (watchdog trip fired, ping resumed) a mechanism for causing trouble 5 minutes later — a "controller side-effects from the re-registration/state-resend persist briefly" theory is speculative, not evidenced. Equally consistent with pure coincidence given this integration's baseline retry rate is already nonzero without any corruption present (e.g. the 2026-09-07 8h36m-armed retry had a fully bus-clean surrounding window).

**Practical takeaway:** no code change. Record as the first candidate data point for the "other signals" search opened in the 2026-09-07 update, but do not treat corruption/watchdog proximity as a working hypothesis yet — one data point with a 5-minute gap is far weaker than the exact-concurrency signature that would make it convincing. Next retry should be checked specifically for whether corruption/watchdog activity is concurrent (same or adjacent minute) rather than just "somewhere in the preceding window," which would be the actual bar for this to become a real lead.

### Update (2026-09-14): a 9h31m-armed disarm retried (2/6) with zero `ff.`/`fe.` corruption or watchdog activity anywhere in the entire surrounding ~27h window — the cleanest data point yet against the corruption-proximity lead

**Source:** the frigate long-term logger, window 2026-09-13 04:14 → 2026-09-14 07:26 UTC (~27h13m). Only one integration-initiated arm/disarm sequence this window.

**Observed facts:**

| Disarm (UTC) | Initiator | Armed since | Duration armed | Retries | `CURRENT_TIME` at disarm | Nearby `ff.`/`fe.`/watchdog activity |
| --- | --- | --- | --- | --- | --- | --- |
| 09-14 04:29:50 | integration (`disarm()`, `ESPHome Keypad`) | 09-13 18:58:41 (physical arm) | ~9h31m | **2/6, ~6.4s (backoff 1s→2s)** | normal, confirmed (glitch-recovery and `Controller time update` broadcasts self-correcting cleanly seconds before and after) | **none** — zero `ff.`/`fe.` occurrences anywhere in the entire 27h13m window, not just nearby (see `protocol_investigations.md`'s 2026-09-14 update) |

**Inference (medium confidence — reinforces the 2026-09-07 revision, further undermines the 2026-09-09 corruption-proximity lead):** this retry occurred in a window with no bus corruption at all, ruling out even loose proximity as an explanation this time — stronger than the 2026-09-09 data point, which at least had a corruption/watchdog episode 5 minutes prior. Combined with the 2026-09-07 8h36m-armed/normal-time retry, this is now the second long-armed, normal-`CURRENT_TIME`, bus-clean retry on record. Neither `CURRENT_TIME`-stuck state, corruption proximity, nor duration-armed alone has held up as a predictor across the full sample; the "other signals" search opened 2026-09-07 has now run through its most obvious candidates without finding one.

**Practical takeaway:** no code change proposed — retry/backoff resolved correctly again (2/6, ~6.4s). Given corruption proximity, `CURRENT_TIME` state, and armed-duration have each been tried and each produced clean counter-examples, a future session should treat this as a lower-priority open question rather than actively hunting for a new candidate signal each time — revisit if a strikingly different case shows up (e.g. a retry that fails to resolve within the existing 6-attempt budget), but routine single retries no longer need dedicated investigation each occurrence.

## Notes

- ARM/STAY/DISARM sequences are simpler than OUTPUT because there's no ACK handshake
Expand All @@ -588,3 +638,4 @@ settle it.
- ARMED_STATE messages (0x11) are published independently regardless of `arm_disarm_state_` (entities always reflect them). As of 2026-08-22, an ARMED_STATE broadcast matching the request's intent resolves the sequence from **any** non-IDLE state (including a retry backoff wait) — see "Retry backoff + intent-matched ARMED_STATE resolution" above. `CODE_ENTER_PENDING` remains the only state that *requires* one to succeed (`ARM_AWAY_PENDING`/`ARM_STAY_PENDING` still resolve on their KEYPAD_COMMAND ack).
- All three keypads receive Command broadcasts during arming/disarming; only the originating keypad controls the sequence
- `disarm()` also guards against calling when already disarmed — it returns early if `is_armed()` is false
- Audible chime count on the official RF remote distinguishes it from other input sources when reading a trace: **one chime on arm, two chimes on disarm** (confirmed by the user, 2026-09-09). This was historically the only way to tell, since an `ARMED_STATE` broadcast considered in isolation looks the same as any other non-integration source — no accompanying keypress/command traffic, unlike physical-keypad or integration-initiated sequences which leave a `Code sequence:`/keypress trail. A direct on-bus signature now also exists: the `0x7C` RF-remote-event frame (see `protocol_wire_format.md`) precedes the resulting `ARMED_STATE` broadcast, typically by ~100–150ms (one case logged in the same millisecond, `0x7C` still ordered first), so a trace no longer has to rely on the chime alone
Loading