Administrator
Published on 2026-09-15 / 18 Visits
0
0

Home Assistant API Automations: A Valid, Unknown, and Stale State Contract

A numeric sensor value is not enough evidence for a Home Assistant automation to act. The system also needs to know whether the value was retrieved successfully, whether it is still fresh, and what should happen when data returns after a failure. A three-state contract makes those hidden assumptions observable: valid, unknown, and stale.

This pattern is useful for prices, weather, tariffs, air quality, irrigation, cloud-connected devices, and any other entity backed by an external API. It prevents two opposite failures: acting on an invalid value and missing an action when a source recovers.

The contract in one table

Quality state Meaning Display policy Control policy
valid The source is available, the value parses, and its age is within the workload's TTL Show the current value Permit the defined action
unknown No trustworthy current observation exists Show the last valid value only with an explicit warning Block consequential actions; alert after a grace period
stale A value exists, but the last trustworthy report is older than its TTL Show the last valid value and age Use a safe fallback or stop; never treat it as current

Home Assistant itself distinguishes unknown from unavailable: its sensor documentation defines the first as a state that is not yet known and the second as an entity that is currently unavailable. The contract groups both under decision quality unknown, while retaining the native reason for diagnostics. A missing entity, parse error, or rejected response can join the same quality state with a different reason.

That preserves useful operational detail without forcing every downstream automation to repeat a growing list of failure strings.

Why a two-state automation fails

A common automation watches a numeric sensor crossing a threshold:

triggers:
  - trigger: numeric_state
    entity_id: sensor.vendor_price
    below: 0.20

This covers the sunny-day path. It leaves several unanswered questions:

  • What if the API returns unknown for 30 seconds?
  • What if the integration becomes unavailable?
  • What if the entity keeps an old numeric value while polling has stopped?
  • What if the API recovers below the threshold, but no new threshold crossing occurs?
  • What if the source flaps between available and unavailable?
  • What if Home Assistant restarts while a for timer is running?

The last two are platform-level details, not theoretical edge cases. Home Assistant's current automation trigger documentation says most entity-targeting triggers do not fire when an entity returns from unknown or unavailable. The state trigger reference also says a for timer resets when Home Assistant restarts or automations reload.

Filtering invalid transitions can prevent false actions, but filtering alone can create a missed-action path. Reliability requires both sides of the state machine: entry into failure and recovery from failure.

Separate value, quality, and authority

The simplest robust design has four objects:

  1. Raw source: the integration-owned entity, such as sensor.vendor_price.
  2. Last valid value: a display and diagnostic copy updated only when the raw value is parseable.
  3. Quality state: one enum that reports valid, unknown, or stale, plus age and reason attributes.
  4. Business automation: the action that consumes the raw value only when quality permits it.

This separation prevents a retained value from silently becoming an authorization token. A last-known temperature is useful on a dashboard. It is unsafe as the only input to a heater controller after the sensor has stopped reporting.

Home Assistant's Template integration supports this pattern directly. Its documentation shows availability, has_value, numeric validation, and a trigger-based template sensor that retains the last valid numeric value.

Here is a minimal last-valid sensor. Replace the entity name and unit with your own:

template:
  - triggers:
      - trigger: state
        entity_id: sensor.vendor_price
        not_to:
          - unknown
          - unavailable
    conditions:
      - condition: template
        value_template: "{{ is_number(states('sensor.vendor_price')) }}"
    sensor:
      - name: "Vendor Price Last Valid"
        unique_id: vendor_price_last_valid
        state: "{{ states('sensor.vendor_price') }}"
        unit_of_measurement: "EUR/kWh"
        attributes:
          value_changed_at: "{{ now().isoformat() }}"

This sensor is intentionally a retained value, not a health signal. value_changed_at describes when this template captured a changed valid value. It does not prove that the upstream API returned fresh data on every poll.

Define freshness by what your timestamp proves

Home Assistant exposes three timestamps on state objects:

  • last_changed: the main state value changed.
  • last_updated: the state or its attributes changed.
  • last_reported: the integration wrote the state to Home Assistant, including an unchanged report.

Current Home Assistant core source includes a state-reported event for an unchanged state write. That makes last_reported the strongest generic Home Assistant-side heartbeat when an integration writes after each successful poll.

It still proves only that Home Assistant received a state write. It does not prove that the provider generated a new observation. An integration can repeatedly report cached data. Prefer these freshness sources in order:

  1. A provider-signed or provider-supplied observation timestamp.
  2. An integration attribute that records a successful upstream fetch.
  3. last_reported, after verifying the integration writes on every successful poll.
  4. last_updated only when a state or attribute change is a valid proxy for a successful fetch.
  5. last_changed only when the business value is expected to change within the TTL.

The distinction matters for stable measurements. A tank level, room temperature, or electricity price may legitimately remain equal across several reports. Home Assistant's state-trigger example for a value that stops changing uses to: null with for. That pattern detects an unchanged value, not a missing report.

Build a centralized quality sensor

The following example assumes the raw entity exists and that last_reported has been validated for the integration. It uses a 10-minute TTL:

template:
  - sensor:
      - name: "Vendor Price Quality"
        unique_id: vendor_price_quality
        state: >
          {% set raw = states('sensor.vendor_price') %}
          {% if raw in ['unknown', 'unavailable'] %}
            unknown
          {% elif not is_number(raw) %}
            unknown
          {% elif (now() - states.sensor.vendor_price.last_reported).total_seconds() > 600 %}
            stale
          {% else %}
            valid
          {% endif %}
        attributes:
          reason: >
            {% set raw = states('sensor.vendor_price') %}
            {% if raw in ['unknown', 'unavailable'] %}
              {{ raw }}
            {% elif not is_number(raw) %}
              parse_error
            {% elif (now() - states.sensor.vendor_price.last_reported).total_seconds() > 600 %}
              ttl_expired
            {% else %}
              ok
            {% endif %}
          source_age_seconds: >
            {{ (now() - states.sensor.vendor_price.last_reported).total_seconds() | int }}
          ttl_seconds: 600
          last_valid_value: "{{ states('sensor.vendor_price_last_valid') }}"

The now() call is deliberate. Home Assistant's date and time templating documentation says templates using now() or utcnow() re-run once per minute, allowing a still-numeric entity to transition into stale as time passes.

Adapt the template before deployment:

  • Use the upstream observation timestamp when one is available.
  • Add range or schema validation, not just numeric parsing.
  • Choose the TTL from the business deadline. A tariff, rain sensor, door state, and air-quality feed need different limits.
  • Handle a missing entity explicitly if your integration can be removed at runtime.
  • Use a trigger-based template with a time pattern if one-minute evaluation is too frequent or too slow.

The quality sensor gives all downstream automations one stable interface. It also gives dashboards and alerts the same vocabulary.

Make recovery a first-class trigger

Suppose an API becomes unknown while the price is high, then returns with a value below the action threshold. An automation that wakes only on a normal numeric transition may not run. Recovery therefore needs its own wake path.

automation:
  - alias: "Run flexible load when price is valid and low"
    mode: single
    triggers:
      - trigger: state
        entity_id: sensor.vendor_price
        to: null
        id: value_changed
      - trigger: state
        entity_id: sensor.vendor_price_quality
        to: valid
        id: recovered
    conditions:
      - condition: state
        entity_id: sensor.vendor_price_quality
        state: valid
      - condition: numeric_state
        entity_id: sensor.vendor_price
        below: 0.20
    actions:
      - action: script.turn_on
        target:
          entity_id: script.set_flexible_load_on

The called script should be idempotent: running it twice must produce the same final device state without duplicate notifications or accumulated side effects. mode: single prevents overlapping runs, but it is not a general deduplication guarantee for two sequential triggers.

For high-impact actions, add three recovery controls:

  • Stabilization: require several successful observations or a bounded stable period before returning to valid.
  • Catch-up window: execute a missed action only if it is still useful. A recovered tariff signal may remain actionable for minutes; an old door-open event should usually not be replayed.
  • Decision key: record the source observation ID, time bucket, or policy version already applied so recovery cannot execute it again.

This converts recovery from an accidental state transition into a defined business event.

Give each state an action budget

Different actions need different failure policies. Use an action matrix instead of one global fallback:

Action class valid unknown stale Recovery
Dashboard display Current value Last valid value plus reason Last valid value plus age Remove warning after stabilization
Notification Normal rule Alert after grace period Alert at TTL, then suppress repeats Send one recovery notice
Reversible comfort control Allow Hold safe current state Apply bounded fallback Re-evaluate current policy
Safety or security control Allow only with all checks Fail closed Fail closed Require fresh evidence; sometimes human approval
Historical logging Record value and quality Record failure reason Record age and TTL Link recovery to the incident

Avoid substituting zero for unknown data unless zero is genuinely a safe and semantically correct value. A default temperature, price, or motion state can authorize the wrong action while hiding the evidence failure.

Add debounce, hysteresis, and alert latching

External APIs fail transiently. Treating one missed poll as an incident creates noise; waiting indefinitely hides outages. Use separate entry and exit rules:

  • Enter unknown after a short grace period or a defined number of consecutive failures.
  • Enter stale exactly when the business TTL expires.
  • Return to valid after one or more successful reports, depending on impact.
  • Send the failure alert once, update its duration, and send one recovery message.
  • Keep control thresholds separate from quality thresholds. Price hysteresis and data freshness solve different problems.

If you implement a grace period with for, remember that its timer resets on Home Assistant restart. For a deadline that must survive restarts, persist the target in an input_datetime helper, as the official template trigger documentation recommends.

Test the contract by injecting failures

A template that renders successfully in Developer Tools has passed a syntax check, not a reliability test. Freeze a small test matrix and record the expected quality transition and action count.

Test Injected condition Expected result
Normal update Fresh valid values above and below the threshold valid; action follows business rule
Native unknown Source becomes unknown Quality becomes unknown; action blocked; reason preserved
Native unavailable Integration or entity becomes unavailable Quality becomes unknown; no default-value action
Frozen numeric value Stop reports while retaining the last number Quality becomes stale at TTL
Recovery Restore a fresh valid value One recovery transition; current policy re-evaluated
Flapping Alternate failure and success rapidly No action or alert storm; stabilization holds
Restart Restart during grace, stale wait, and recovery Persistent deadlines remain correct or reset by documented design
Missing entity Remove or rename the source Quality fails closed with a diagnostic reason

For each row, inspect automation traces and history. Count business actions, notifications, quality transitions, and time to recovery. The acceptance condition is an observable state path, not a reassuring final dashboard.

Start with the minimum contract

Do not build a general reliability framework before one real API works end to end. Start with one raw source, one quality sensor, one TTL, one idempotent action, one alert, and the test matrix. Add integration-specific timestamps, consecutive-failure counters, persisted deadlines, or fleet-wide aggregation only after real failures justify them.

This keeps the contract easy to operate. It also makes failure data useful: every manual intervention can be classified as a missing state, missing transition, wrong TTL, duplicate action, or inadequate fallback.

For adjacent architecture patterns, see Local Vision, Cloud Reasoning and AI Agent Policy Compliance: From Context to Commit Gates.

Frequently asked questions

What is the difference between unknown and unavailable in Home Assistant?

unknown means the entity's state is not yet known. unavailable means the entity is currently unavailable. A business quality layer can map both to unknown while retaining the native reason for diagnostics and recovery policy.

Should an automation ignore transitions from unknown or unavailable?

Ignoring them can prevent false triggers, but it can also hide recovery. Filter invalid business actions and add a separate recovery trigger that re-evaluates the current rule with fresh evidence.

How do I detect a stale Home Assistant sensor?

Compare the current time with a timestamp that actually represents successful data receipt. Prefer an upstream observation time. Use last_reported only after verifying the integration writes on every successful poll. A value that has not changed is not automatically stale.

Should I keep the last valid sensor value?

Yes for display, diagnostics, and some bounded fallbacks. Store quality separately and require valid for consequential control. A retained value without age can look current after the source has failed.

Does a state trigger with for survive a Home Assistant restart?

No. The official state trigger reference says the timer resets on restart or automation reload. Store a persistent deadline in a helper when the timing contract must survive restarts.

Is availability enough?

availability prevents a derived entity from presenting invalid input as a normal value. It does not define staleness, recovery replay, action authority, alert suppression, or persisted deadlines. Use it as one part of the contract.

References


Comment