2018-09-27 16:50:00 -07:00
|
|
|
/*
|
|
|
|
|
* Copyright (c) 2018 Intel Corporation
|
|
|
|
|
*
|
|
|
|
|
* SPDX-License-Identifier: Apache-2.0
|
|
|
|
|
*/
|
2019-10-25 00:08:21 +09:00
|
|
|
|
sys: util: move lowercase min/max/clamp to a new minmax.h
Since commit 37717b229f51 ("sys: util: rename Z_MIN Z_MAX Z_CLAMP to min
max and clamp"), <zephyr/sys/util.h> unconditionally defines function-
like macros named `min`, `max`, and `clamp` in the global namespace (in
C mode). util.h gets pulled in transitively by very broad headers,
including the POSIX layer's <pthread.h>, so any third-party C code that
uses these names as ordinary identifiers (e.g. XNNPACK's static `clamp`
helper and its public `clamp` struct field) fails to build as soon as
<pthread.h> is included.
Following the approach used by Linux, move the lowercase `min`, `max`,
`min3`, `max3`, and `clamp` macros (and their helpers) into a new
<zephyr/sys/minmax.h> header that has to be included explicitly by
source files that want them. util.h keeps the uppercase MIN/MAX/CLAMP,
so most code is unaffected; only the (much smaller) set of files that
actually use the lowercase variants needs to pick up the new include.
Fixes #107853.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-15 12:13:48 -04:00
|
|
|
#include <zephyr/sys/minmax.h>
|
2022-05-06 11:04:23 +02:00
|
|
|
#include <zephyr/kernel.h>
|
|
|
|
|
#include <zephyr/spinlock.h>
|
2018-09-27 16:50:00 -07:00
|
|
|
#include <ksched.h>
|
2023-08-29 19:32:46 +00:00
|
|
|
#include <timeout_q.h>
|
2023-09-26 22:46:01 +00:00
|
|
|
#include <zephyr/internal/syscall_handler.h>
|
2022-05-06 11:04:23 +02:00
|
|
|
#include <zephyr/drivers/timer/system_timer.h>
|
2026-06-26 10:43:03 +02:00
|
|
|
#include <zephyr/sys/clock.h>
|
2025-12-09 10:09:23 -08:00
|
|
|
#include <zephyr/llext/symbol.h>
|
2018-09-27 16:50:00 -07:00
|
|
|
|
2026-04-06 17:11:00 -04:00
|
|
|
#include <timeslicing.h>
|
|
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
/* Absolute tick counter: ticks since boot. */
|
2020-05-27 11:26:57 -05:00
|
|
|
static uint64_t curr_tick;
|
2018-09-27 16:50:00 -07:00
|
|
|
|
2025-01-09 13:50:38 -08:00
|
|
|
/*
|
|
|
|
|
* The timeout code shall take no locks other than its own (timeout_lock), nor
|
kernel: timeout: add timer-wheel backend
Add a hierarchical timer wheel as a third selectable timeout backend.
Timeouts are bucketed by expiry distance: one list per tick for the
next 32 ticks ("soon"), one list per 32-tick band for the next ~1024
("later"), and a sorted overflow list ("distant"). Insertion and
removal are O(1) for the common near-future case; every 32 ticks the
announce path sifts the current "later" band into "soon" and refills
it from "distant". This scales well when many short-lived timeouts are
pending.
Unlike the dlist and min-heap backends, the wheel does not fit the
generic next_gap/advance/pop_due announce primitives: its per-tick
advance is a bitmap scan that jumps over empty ticks, and its sift is
a time-driven event tied to no single timeout. Rather than contort
that (already subtle) state machine, the wheel uses the backend-owned
announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and
implements z_timeout_q_announce(), which sys_clock_announce_locked()
calls in place of the generic loop.
Like the other backends the wheel is a single implementation header
(kernel/timeout_wheel.h) included only by timeout.c, so its state, its
operations and that announce loop reach the shared state (curr_tick,
announce_remaining, inflight_timeout) directly; nothing extra needs
exposing. The SMP re-entry guard, announcing_cpu and the reprogram
remain in timeout.c. The wheel fires handlers through the same
inflight_timeout dance, so the post-#109977 abort/in-flight
synchronization works unchanged, and it carries no per-node
ANNOUNCING/ABORTED sentinels.
struct _timeout grows a wheel-only flags field (which wheel tier a
timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist
and heap builds are unaffected.
The backend is EXPERIMENTAL. Two known limitations, inherent to the
wheel algorithm (not the abstraction):
- No same-tick firing-order guarantee (sifted timeouts are
prepended), so the timeout_order test does not apply.
- next_timeout() never exceeds 32 ticks because a sift is always
pending, so the wheel wakes a tickless-idle CPU at least every 32
ticks. This fails tests that assert zero spurious idle wakeups
(tests/kernel/context cpu_idle / timer_interrupts) and is a power
regression versus the dlist and heap backends.
Verified the timeout-functional suites (timer_api, timeout,
timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86,
x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected.
The timer-wheel data structure, bucketing scheme and sift algorithm
are the work of Peter Mitsis (PR #108339), re-homed here behind the
timeout backend interface.
Co-authored-by: Peter Mitsis <peter.mitsis@intel.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
|
|
|
* shall it call any other subsystem while holding this lock. Code outside this
|
|
|
|
|
* file takes it through sys_clock_lock()/sys_clock_unlock().
|
2025-01-09 13:50:38 -08:00
|
|
|
*/
|
2018-09-27 16:50:00 -07:00
|
|
|
static struct k_spinlock timeout_lock;
|
|
|
|
|
|
2023-06-28 23:17:45 +02:00
|
|
|
/* Ticks left to process in the currently-executing sys_clock_announce() */
|
2026-06-10 14:42:13 -04:00
|
|
|
static uint32_t announce_remaining;
|
2018-09-27 16:50:00 -07:00
|
|
|
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
/* CPU id currently inside sys_clock_announce_locked()'s firing loop, or -1
|
|
|
|
|
* when no CPU is. The SMP early-return below ensures at most one CPU is in
|
|
|
|
|
* the loop at a time, so a single int suffices. Used by code that needs to
|
|
|
|
|
* know whether *this* CPU is at a tick edge (announcing) versus somewhere
|
|
|
|
|
* within a tick (any other context, including running on a CPU while a
|
|
|
|
|
* different CPU is the announcer).
|
|
|
|
|
*/
|
|
|
|
|
static int announcing_cpu = -1;
|
|
|
|
|
|
|
|
|
|
static inline bool this_cpu_announcing(void)
|
|
|
|
|
{
|
|
|
|
|
return announcing_cpu == CPU_ID;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
static inline bool any_cpu_announcing(void)
|
|
|
|
|
{
|
|
|
|
|
return announcing_cpu != -1;
|
|
|
|
|
}
|
|
|
|
|
|
2026-05-27 13:18:00 -04:00
|
|
|
/* Timeout whose handler is currently being dispatched, or NULL when no
|
kernel: timeout: tag inflight_timeout with a superseded bit
When z_try_abort_timeout() finds the target timeout already popped
from the queue and in flight (its handler dispatching), the abort
cannot remove anything from the queue. The handler is about to run
(or is running) the timeout's callback; the aborter needs a way to
tell it "you were aborted, skip your side effects". main carried that
signal in the per-timeout dticks=ABORTED sentinel, which any CPU
could write under timeout_lock and the handler checked at entry via
z_is_timeout_handler_canceled(). Later commits in this series remove
the dticks-cancel mechanism, so the signal needs a new home.
Encode it in the low bit of the file-local inflight_timeout pointer
(struct _timeout is pointer-aligned, so bit 0 is free):
inflight_timeout == NULL no handler in flight
inflight_timeout == t handler in flight, not superseded
inflight_timeout == t | 1 handler in flight, superseded
z_try_abort_timeout() sets the bit whenever the target is the
in-flight timeout -- on the announcing CPU (same-CPU IRQ that
preempted the dispatch, or a stop after a re-arm) and on another CPU
racing the handler.
A handler with non-idempotent side effects (currently only k_timer's
expiry_fn) checks z_timeout_inflight_superseded() at entry and bails
if set. Idempotent handlers (z_thread_timeout, work, poll, ...)
tolerate the race and don't need the check.
The bit is a best-effort signal, not a barrier: an aborter that sets
it after the handler has passed its check has no effect (the handler
already committed), exactly as the dticks sentinel behaved on main.
The -EAGAIN return for the cross-CPU case is a separate mechanism,
used only by callers that must wait for the handler to fully complete
(e.g. before freeing the timeout's storage); they spin, while
best-effort callers ignore -EAGAIN and rely on the bit.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
|
|
|
* handler is in flight. The announcing CPU sets the pointer before
|
|
|
|
|
* calling the handler and clears it afterwards; any aborter may set the
|
|
|
|
|
* low "superseded" bit (struct _timeout is pointer-aligned so bit 0 is
|
|
|
|
|
* free). Accessors below mask the bit.
|
2026-05-27 13:18:00 -04:00
|
|
|
*/
|
|
|
|
|
static struct _timeout *inflight_timeout;
|
|
|
|
|
|
kernel: timeout: tag inflight_timeout with a superseded bit
When z_try_abort_timeout() finds the target timeout already popped
from the queue and in flight (its handler dispatching), the abort
cannot remove anything from the queue. The handler is about to run
(or is running) the timeout's callback; the aborter needs a way to
tell it "you were aborted, skip your side effects". main carried that
signal in the per-timeout dticks=ABORTED sentinel, which any CPU
could write under timeout_lock and the handler checked at entry via
z_is_timeout_handler_canceled(). Later commits in this series remove
the dticks-cancel mechanism, so the signal needs a new home.
Encode it in the low bit of the file-local inflight_timeout pointer
(struct _timeout is pointer-aligned, so bit 0 is free):
inflight_timeout == NULL no handler in flight
inflight_timeout == t handler in flight, not superseded
inflight_timeout == t | 1 handler in flight, superseded
z_try_abort_timeout() sets the bit whenever the target is the
in-flight timeout -- on the announcing CPU (same-CPU IRQ that
preempted the dispatch, or a stop after a re-arm) and on another CPU
racing the handler.
A handler with non-idempotent side effects (currently only k_timer's
expiry_fn) checks z_timeout_inflight_superseded() at entry and bails
if set. Idempotent handlers (z_thread_timeout, work, poll, ...)
tolerate the race and don't need the check.
The bit is a best-effort signal, not a barrier: an aborter that sets
it after the handler has passed its check has no effect (the handler
already committed), exactly as the dticks sentinel behaved on main.
The -EAGAIN return for the cross-CPU case is a separate mechanism,
used only by callers that must wait for the handler to fully complete
(e.g. before freeing the timeout's storage); they spin, while
best-effort callers ignore -EAGAIN and rely on the bit.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
|
|
|
#define INFLIGHT_SUPERSEDED_BIT 1UL
|
|
|
|
|
|
|
|
|
|
static inline struct _timeout *inflight_ptr(void)
|
|
|
|
|
{
|
|
|
|
|
return (struct _timeout *)((uintptr_t)inflight_timeout & ~INFLIGHT_SUPERSEDED_BIT);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
static inline void inflight_mark_superseded(void)
|
|
|
|
|
{
|
|
|
|
|
inflight_timeout = (struct _timeout *)((uintptr_t)inflight_timeout |
|
|
|
|
|
INFLIGHT_SUPERSEDED_BIT);
|
|
|
|
|
}
|
|
|
|
|
|
2026-06-10 14:42:13 -04:00
|
|
|
static uint32_t elapsed(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
/*
|
|
|
|
|
* While *this* CPU is executing sys_clock_announce_locked()'s firing
|
|
|
|
|
* loop, new relative timeouts scheduled from a callback (or from a
|
|
|
|
|
* higher-priority ISR that preempted one) are anchored to the currently
|
|
|
|
|
* firing tick (curr_tick), so we report 0.
|
2023-06-28 23:17:45 +02:00
|
|
|
*
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
* On any other CPU the picture is different: we are not at a tick edge,
|
|
|
|
|
* and curr_tick is partway through being advanced by the announcing
|
|
|
|
|
* CPU's loop. The invariant we want to preserve is
|
2023-06-28 23:17:45 +02:00
|
|
|
*
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
* T_real - curr_tick = announce_remaining + sys_clock_elapsed()
|
2023-06-28 23:17:45 +02:00
|
|
|
*
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
* because the driver bumped its internal announced-cycle baseline to
|
|
|
|
|
* (curr_tick_initial + N) * CYC_PER_TICK at ISR entry, while the kernel
|
|
|
|
|
* has only advanced curr_tick by the K ticks processed so far -- so
|
|
|
|
|
* announce_remaining (= N - K) is the residual that must be added to
|
|
|
|
|
* sys_clock_elapsed() to get the real-time delta from curr_tick. This
|
|
|
|
|
* keeps sys_clock_tick_get() monotonic across the announce window.
|
2023-06-28 23:17:45 +02:00
|
|
|
*/
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
if (this_cpu_announcing()) {
|
|
|
|
|
return 0U;
|
|
|
|
|
}
|
2026-05-25 15:12:26 -04:00
|
|
|
return sys_clock_elapsed() +
|
|
|
|
|
(IS_ENABLED(CONFIG_SMP) ? announce_remaining : 0);
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
/*
|
|
|
|
|
* Backend queue implementation. The selected backend's queue instance and its
|
|
|
|
|
* z_timeout_q_*() operations are private to this file: the backend is a header
|
|
|
|
|
* included only here, after the shared state above, so those operations can
|
|
|
|
|
* reach curr_tick / announce_remaining / inflight_timeout / timeout_lock
|
|
|
|
|
* directly. The per-node helpers (z_init_timeout / z_is_inactive_timeout) are
|
|
|
|
|
* tree-wide and live in timeout_q.h.
|
|
|
|
|
*/
|
kernel: timeout: add min-heap backend
Add a binary min-heap as a selectable timeout backend alongside the
sorted delta list, plugging into the backend abstraction rather than
forking timeout.c. Pending timeouts are kept in a min-heap keyed on
absolute expiry tick; insertion and arbitrary removal are O(log n),
which scales better than the delta list's O(n) insertion when many
timeouts are pending.
The backend is a new kernel/timeout_minheap.h: the heap instance (a
min_heap_ref from the previous commit), its comparator and the
z_timeout_q_*() operations, included only by timeout.c (after curr_tick)
so the whole backend stays private to that translation unit. struct
_timeout gains the abs_ticks + heap_handle representation, selected by
Kconfig (the delta list's node + dticks is the #else), with the common
fn pointer kept as a shared trailing member; the per-node helpers in
timeout_q.h grow a matching min-heap variant. The kernel shell thread
dump prints the backend's raw scheduling field (dticks or abs_ticks)
under #ifdef.
Because the shared z_add_timeout() already applies the post-#107452
conditional tick round-up, the heap inherits it: the backend's insert
simply stores abs_ticks = curr_tick + dticks. Likewise in-flight
handler synchronization (PR #109977) lives in timeout.c, so the heap
node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just
heap_handle.idx == 0.
The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks
need 64-bit precision), and uses a fixed-capacity heap
(CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error.
The min-heap algorithm, struct fields, Kconfig and capacity model are
derived from Sayooj K Karun's min-heap timeout subsystem (#106013),
reworked here to fit the pluggable backend and the current in-tree
in-flight-handler synchronization.
Co-authored-by: Sayooj K Karun <sayooj@aerlync.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
|
|
|
#if defined(CONFIG_TIMEOUT_BACKEND_MINHEAP)
|
|
|
|
|
#include "timeout_minheap.h"
|
kernel: timeout: add timer-wheel backend
Add a hierarchical timer wheel as a third selectable timeout backend.
Timeouts are bucketed by expiry distance: one list per tick for the
next 32 ticks ("soon"), one list per 32-tick band for the next ~1024
("later"), and a sorted overflow list ("distant"). Insertion and
removal are O(1) for the common near-future case; every 32 ticks the
announce path sifts the current "later" band into "soon" and refills
it from "distant". This scales well when many short-lived timeouts are
pending.
Unlike the dlist and min-heap backends, the wheel does not fit the
generic next_gap/advance/pop_due announce primitives: its per-tick
advance is a bitmap scan that jumps over empty ticks, and its sift is
a time-driven event tied to no single timeout. Rather than contort
that (already subtle) state machine, the wheel uses the backend-owned
announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and
implements z_timeout_q_announce(), which sys_clock_announce_locked()
calls in place of the generic loop.
Like the other backends the wheel is a single implementation header
(kernel/timeout_wheel.h) included only by timeout.c, so its state, its
operations and that announce loop reach the shared state (curr_tick,
announce_remaining, inflight_timeout) directly; nothing extra needs
exposing. The SMP re-entry guard, announcing_cpu and the reprogram
remain in timeout.c. The wheel fires handlers through the same
inflight_timeout dance, so the post-#109977 abort/in-flight
synchronization works unchanged, and it carries no per-node
ANNOUNCING/ABORTED sentinels.
struct _timeout grows a wheel-only flags field (which wheel tier a
timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist
and heap builds are unaffected.
The backend is EXPERIMENTAL. Two known limitations, inherent to the
wheel algorithm (not the abstraction):
- No same-tick firing-order guarantee (sifted timeouts are
prepended), so the timeout_order test does not apply.
- next_timeout() never exceeds 32 ticks because a sift is always
pending, so the wheel wakes a tickless-idle CPU at least every 32
ticks. This fails tests that assert zero spurious idle wakeups
(tests/kernel/context cpu_idle / timer_interrupts) and is a power
regression versus the dlist and heap backends.
Verified the timeout-functional suites (timer_api, timeout,
timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86,
x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected.
The timer-wheel data structure, bucketing scheme and sift algorithm
are the work of Peter Mitsis (PR #108339), re-homed here behind the
timeout backend interface.
Co-authored-by: Peter Mitsis <peter.mitsis@intel.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
|
|
|
#elif defined(CONFIG_TIMEOUT_BACKEND_WHEEL)
|
|
|
|
|
#include "timeout_wheel.h"
|
kernel: timeout: add bucketed delta-list backend
Add a single-level bucketed delta list as a fourth selectable timeout
backend: a simpler relative of the timer wheel that keeps the wheel's
O(1) near-future insertion without its second tier, sift/defer
machinery, or tickless-idle penalty.
Timeouts expiring within the next CONFIG_TIMEOUT_BUCKET_LISTS ticks go
into a per-tick bucket list (O(1) insert and remove), tracked by an
occupancy bitmap so "ticks until next expiry" is O(1). Everything
beyond goes into one overflow list sorted by absolute expiry (O(n)
insert, O(1) remove, no delta fix-up). As curr_tick crosses the bucket
window, overflow entries that have come within range migrate into
buckets.
Like the other backends it is a single implementation header
(kernel/timeout_bucket.h) included only by timeout.c. Two properties
distinguish it from the other non-default backends, both confirmed by
running the full kernel timer/context/common suites with the backend
forced on (qemu_x86, x86_64, cortex_a53 SMP, riscv64):
- It fits the generic next_gap/advance/pop_due announce loop -- no
backend-owned announce. Migration is on demand against the actual
next event, so an idle CPU with a distant timeout sleeps straight
to it: it passes tests/kernel/context cpu_idle / timer_interrupts,
which the wheel fails (the wheel wakes at least every 32 ticks).
- Same-tick firing order stays FIFO (bucket and overflow inserts
append; migration preserves order), so tests/kernel/common's
timeout_order passes, unlike the min-heap.
The per-node representation is the delta list's node + dticks with no
extra field: a bucket entry stores its bucket index (< BUCKET_LISTS) in
dticks, an overflow entry its absolute expiry (>= BUCKET_LISTS); the
ranges never overlap, so it shares the dlist per-node helpers. Absolute
expiry needs 64-bit ticks, so the backend depends on TIMEOUT_64BIT. It
inherits the shared z_add_timeout round-up and the inflight_timeout
synchronization like the other backends. EXPERIMENTAL.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 22:00:48 -04:00
|
|
|
#elif defined(CONFIG_TIMEOUT_BACKEND_BUCKET)
|
|
|
|
|
#include "timeout_bucket.h"
|
2026-08-25 21:39:08 +05:30
|
|
|
#elif defined(CONFIG_TIMEOUT_BACKEND_SKIPLIST)
|
|
|
|
|
#include "timeout_skiplist.h"
|
kernel: timeout: add min-heap backend
Add a binary min-heap as a selectable timeout backend alongside the
sorted delta list, plugging into the backend abstraction rather than
forking timeout.c. Pending timeouts are kept in a min-heap keyed on
absolute expiry tick; insertion and arbitrary removal are O(log n),
which scales better than the delta list's O(n) insertion when many
timeouts are pending.
The backend is a new kernel/timeout_minheap.h: the heap instance (a
min_heap_ref from the previous commit), its comparator and the
z_timeout_q_*() operations, included only by timeout.c (after curr_tick)
so the whole backend stays private to that translation unit. struct
_timeout gains the abs_ticks + heap_handle representation, selected by
Kconfig (the delta list's node + dticks is the #else), with the common
fn pointer kept as a shared trailing member; the per-node helpers in
timeout_q.h grow a matching min-heap variant. The kernel shell thread
dump prints the backend's raw scheduling field (dticks or abs_ticks)
under #ifdef.
Because the shared z_add_timeout() already applies the post-#107452
conditional tick round-up, the heap inherits it: the backend's insert
simply stores abs_ticks = curr_tick + dticks. Likewise in-flight
handler synchronization (PR #109977) lives in timeout.c, so the heap
node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just
heap_handle.idx == 0.
The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks
need 64-bit precision), and uses a fixed-capacity heap
(CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error.
The min-heap algorithm, struct fields, Kconfig and capacity model are
derived from Sayooj K Karun's min-heap timeout subsystem (#106013),
reworked here to fit the pluggable backend and the current in-tree
in-flight-handler synchronization.
Co-authored-by: Sayooj K Karun <sayooj@aerlync.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
|
|
|
#else /* CONFIG_TIMEOUT_BACKEND_DLIST */
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
#include "timeout_list.h"
|
kernel: timeout: add min-heap backend
Add a binary min-heap as a selectable timeout backend alongside the
sorted delta list, plugging into the backend abstraction rather than
forking timeout.c. Pending timeouts are kept in a min-heap keyed on
absolute expiry tick; insertion and arbitrary removal are O(log n),
which scales better than the delta list's O(n) insertion when many
timeouts are pending.
The backend is a new kernel/timeout_minheap.h: the heap instance (a
min_heap_ref from the previous commit), its comparator and the
z_timeout_q_*() operations, included only by timeout.c (after curr_tick)
so the whole backend stays private to that translation unit. struct
_timeout gains the abs_ticks + heap_handle representation, selected by
Kconfig (the delta list's node + dticks is the #else), with the common
fn pointer kept as a shared trailing member; the per-node helpers in
timeout_q.h grow a matching min-heap variant. The kernel shell thread
dump prints the backend's raw scheduling field (dticks or abs_ticks)
under #ifdef.
Because the shared z_add_timeout() already applies the post-#107452
conditional tick round-up, the heap inherits it: the backend's insert
simply stores abs_ticks = curr_tick + dticks. Likewise in-flight
handler synchronization (PR #109977) lives in timeout.c, so the heap
node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just
heap_handle.idx == 0.
The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks
need 64-bit precision), and uses a fixed-capacity heap
(CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error.
The min-heap algorithm, struct fields, Kconfig and capacity model are
derived from Sayooj K Karun's min-heap timeout subsystem (#106013),
reworked here to fit the pluggable backend and the current in-tree
in-flight-handler synchronization.
Co-authored-by: Sayooj K Karun <sayooj@aerlync.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
|
|
|
#endif
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
|
|
|
|
|
/*
|
|
|
|
|
* Ticks the driver may wait before the next sys_clock_announce(). The
|
|
|
|
|
* announce-range cap lives here, in the backend-independent front end, so no
|
|
|
|
|
* backend has to reproduce it: the backend only reports the delta to its
|
|
|
|
|
* earliest pending timeout via z_timeout_q_next_expiry().
|
|
|
|
|
*/
|
2026-06-10 17:03:30 -04:00
|
|
|
static uint32_t next_timeout(uint32_t ticks_elapsed)
|
2019-01-16 08:54:38 -08:00
|
|
|
{
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
k_ticks_t next = z_timeout_q_next_expiry();
|
2026-06-10 17:03:30 -04:00
|
|
|
uint32_t dticks;
|
2022-02-08 15:28:09 -08:00
|
|
|
|
2026-06-10 17:03:30 -04:00
|
|
|
/*
|
|
|
|
|
* sys_clock_announce() reports the ticks elapsed since the previous
|
|
|
|
|
* announce, so the gap between two announces must not exceed
|
|
|
|
|
* SYS_CLOCK_MAX_WAIT for the announced count to fit. The budget left
|
|
|
|
|
* from now on is SYS_CLOCK_MAX_WAIT - ticks_elapsed; if it is already
|
|
|
|
|
* spent ask for an announce right away.
|
|
|
|
|
*/
|
|
|
|
|
if (ticks_elapsed >= SYS_CLOCK_MAX_WAIT) {
|
|
|
|
|
return 0;
|
2022-02-08 15:28:09 -08:00
|
|
|
}
|
2019-01-16 08:54:38 -08:00
|
|
|
|
2026-06-10 17:03:30 -04:00
|
|
|
/*
|
2026-07-30 18:01:00 -04:00
|
|
|
* No deadline, or one further out than can be scheduled in a single
|
|
|
|
|
* step: wait the capped budget and re-evaluate at the next announce.
|
|
|
|
|
* Testing next (which may be 64-bit) against the cap keeps the
|
|
|
|
|
* remaining arithmetic in 32 bits. The empty case still returns the
|
|
|
|
|
* budget so a driver keeps waking to maintain uptime.
|
2026-06-10 17:03:30 -04:00
|
|
|
*/
|
2026-07-30 18:01:00 -04:00
|
|
|
if (next == K_TICKS_FOREVER || next >= SYS_CLOCK_MAX_WAIT) {
|
2026-06-10 17:03:30 -04:00
|
|
|
return SYS_CLOCK_MAX_WAIT - ticks_elapsed;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
/* Otherwise wait until the timeout, relative to now (0 if due). */
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
dticks = (uint32_t)next;
|
2026-06-10 17:03:30 -04:00
|
|
|
|
|
|
|
|
return (dticks > ticks_elapsed) ? (dticks - ticks_elapsed) : 0;
|
2019-01-16 08:54:38 -08:00
|
|
|
}
|
|
|
|
|
|
2026-07-30 18:01:00 -04:00
|
|
|
/*
|
|
|
|
|
* Reprogram the timer for the next pending timeout, or, when nothing is
|
|
|
|
|
* pending and CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE tolerates a drifting uptime, tell
|
|
|
|
|
* the driver its clock is unused so it may stop. Only meaningful where the
|
|
|
|
|
* timeout queue may have just drained (abort, end of announce); the add path
|
|
|
|
|
* always has a pending timeout and calls sys_clock_set_timeout() directly.
|
|
|
|
|
*/
|
|
|
|
|
static void reprogram_next(uint32_t ticks_elapsed)
|
|
|
|
|
{
|
|
|
|
|
if (IS_ENABLED(CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE) &&
|
|
|
|
|
z_timeout_q_next_expiry() == K_TICKS_FOREVER) {
|
|
|
|
|
sys_clock_no_timeout();
|
|
|
|
|
} else {
|
|
|
|
|
sys_clock_set_timeout(next_timeout(ticks_elapsed), false);
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
2025-04-04 09:33:03 +02:00
|
|
|
k_ticks_t z_add_timeout(struct _timeout *to, _timeout_func_t fn, k_timeout_t timeout)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
2025-04-04 09:33:03 +02:00
|
|
|
k_ticks_t ticks = 0;
|
|
|
|
|
|
2020-04-21 11:07:07 -07:00
|
|
|
if (K_TIMEOUT_EQ(timeout, K_FOREVER)) {
|
2025-04-04 09:33:03 +02:00
|
|
|
return 0;
|
2020-04-21 11:07:07 -07:00
|
|
|
}
|
|
|
|
|
|
2020-12-07 13:15:42 -05:00
|
|
|
#ifdef CONFIG_KERNEL_COHERENCE
|
2025-10-29 11:52:18 -07:00
|
|
|
__ASSERT_NO_MSG(sys_cache_is_mem_coherent(to));
|
2024-03-08 12:00:10 +01:00
|
|
|
#endif /* CONFIG_KERNEL_COHERENCE */
|
kernel: Add cache coherence management framework
Zephyr SMP kernels need to be able to run on architectures with
incoherent caches. Naive implementation of synchronization on such
architectures requires extensive cache flushing (e.g. flush+invalidate
everything on every spin lock operation, flush on every unlock!) and
is a performance problem.
Instead, many of these systems will have access to separate "coherent"
(usually uncached) and "incoherent" regions of memory. Where this is
available, place all writable data sections by default into the
coherent region. An "__incoherent" attribute flag is defined for data
regions that are known to be CPU-local and which should use the cache.
By default, this is used for stack memory.
Stack memory will be incoherent by default, as by definition it is
local to its current thread. This requires special cache management
on context switch, so an arch API has been added for that.
Also, when enabled, add assertions to strategic places to ensure that
shared kernel data is indeed coherent. We check thread objects, the
_kernel struct, waitq's, timeouts and spinlocks. In practice almost
all kernel synchronization is built on top of these structures, and
any shared data structs will contain at least one of them.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2020-05-13 15:34:04 +00:00
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
__ASSERT(z_is_inactive_timeout(to), "");
|
2018-09-27 16:50:00 -07:00
|
|
|
to->fn = fn;
|
|
|
|
|
|
2023-07-07 09:12:38 +02:00
|
|
|
K_SPINLOCK(&timeout_lock) {
|
2026-09-06 00:39:45 +05:30
|
|
|
uint32_t ticks_elapsed = 0;
|
2025-03-31 13:25:03 +02:00
|
|
|
bool has_elapsed = false;
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
k_ticks_t dticks;
|
2018-09-27 16:50:00 -07:00
|
|
|
|
2025-02-25 16:02:46 -08:00
|
|
|
if (Z_IS_TIMEOUT_RELATIVE(timeout)) {
|
2025-03-31 13:25:03 +02:00
|
|
|
ticks_elapsed = elapsed();
|
|
|
|
|
has_elapsed = true;
|
kernel: timeout: make z_add_timeout round-up conditional on announce
z_add_timeout() has always added one tick to the incoming tick count,
as a conservative round-up so that a request issued partway through a
tick still waits for "at least N full ticks" before the fire. This
round-up is correct in the general case, but *wrong* when the call
happens from within sys_clock_announce_locked() -- i.e. from a timer
expiration callback running at the tick-processing boundary where
elapsed() already returns 0. In that context there is no fractional
tick to compensate for, and the round-up simply makes every scheduled
timeout one tick late.
Two in-tree callers were already compensating for this caller-side:
* The k_timer periodic reschedule path in z_timer_expiration_handler
always runs from inside sys_clock_announce_locked() and subtracted
1 from the period before calling z_add_timeout(). Under
CONFIG_TIMEOUT_64BIT, the same path additionally added +1 inside
K_TIMEOUT_ABS_TICKS() to undo a related round-down.
* z_time_slice_reset() armed the slice timer with K_TICKS(slice_size
- 1) so the resulting fire would land at exactly slice_size ticks.
This one is reachable from both thread context (the +1 cancels the
-1) and from announce context via update_cache() during a
ready-thread wakeup (the +1 isn't applied, leaving the slice short
by one tick). The latter is what actually trips
tests/kernel/tickless/tickless_concept on every platform once the
conditional below lands without dropping these workarounds.
All three are symptoms of the same root cause.
Handle it at the source: make the +1 conditional on announce_remaining
== 0. When scheduling from the timer ISR, we are already at a tick
boundary by construction, so no round-up is needed and periodic timers
now reschedule at exact period intervals without any caller-side
compensation. Drop the -1 in z_timer_expiration_handler's period
path, the +1 in its 64-bit absolute-reschedule companion, and the -1
in z_time_slice_reset(), since all three existed solely to paper over
this mismatch.
This change only affects timeouts scheduled from announce context
(periodic k_timers rescheduling themselves, callbacks starting new
timers, slice-timer rearm during ready-thread wakeup). All other
callers -- k_sleep(), z_abort_timeout(), initial k_timer_start() from
a thread, k_sched_time_slice_set() -- continue to use the +1 round-up.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-17 00:36:30 -04:00
|
|
|
/*
|
|
|
|
|
* In the general case, "now" may be anywhere within
|
|
|
|
|
* the current tick. Rounding up by one tick guarantees
|
|
|
|
|
* "at least N ticks" semantics -- otherwise a request
|
|
|
|
|
* made partway through a tick would fire on the next
|
|
|
|
|
* tick edge, yielding less than N full ticks.
|
|
|
|
|
*
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
* The one moment we know *this* CPU is at a tick edge
|
|
|
|
|
* is while it is processing timeouts inside its own
|
|
|
|
|
* sys_clock_announce_locked() loop. Periodic timers
|
|
|
|
|
* rely on this when rescheduling themselves from the
|
|
|
|
|
* timer ISR: the round-up would otherwise accumulate
|
|
|
|
|
* and make every period one tick late. The check is
|
|
|
|
|
* per-CPU: a thread on another CPU running while we
|
|
|
|
|
* announce is *not* at a tick edge and still needs
|
|
|
|
|
* the round-up.
|
kernel: timeout: make z_add_timeout round-up conditional on announce
z_add_timeout() has always added one tick to the incoming tick count,
as a conservative round-up so that a request issued partway through a
tick still waits for "at least N full ticks" before the fire. This
round-up is correct in the general case, but *wrong* when the call
happens from within sys_clock_announce_locked() -- i.e. from a timer
expiration callback running at the tick-processing boundary where
elapsed() already returns 0. In that context there is no fractional
tick to compensate for, and the round-up simply makes every scheduled
timeout one tick late.
Two in-tree callers were already compensating for this caller-side:
* The k_timer periodic reschedule path in z_timer_expiration_handler
always runs from inside sys_clock_announce_locked() and subtracted
1 from the period before calling z_add_timeout(). Under
CONFIG_TIMEOUT_64BIT, the same path additionally added +1 inside
K_TIMEOUT_ABS_TICKS() to undo a related round-down.
* z_time_slice_reset() armed the slice timer with K_TICKS(slice_size
- 1) so the resulting fire would land at exactly slice_size ticks.
This one is reachable from both thread context (the +1 cancels the
-1) and from announce context via update_cache() during a
ready-thread wakeup (the +1 isn't applied, leaving the slice short
by one tick). The latter is what actually trips
tests/kernel/tickless/tickless_concept on every platform once the
conditional below lands without dropping these workarounds.
All three are symptoms of the same root cause.
Handle it at the source: make the +1 conditional on announce_remaining
== 0. When scheduling from the timer ISR, we are already at a tick
boundary by construction, so no round-up is needed and periodic timers
now reschedule at exact period intervals without any caller-side
compensation. Drop the -1 in z_timer_expiration_handler's period
path, the +1 in its 64-bit absolute-reschedule companion, and the -1
in z_time_slice_reset(), since all three existed solely to paper over
this mismatch.
This change only affects timeouts scheduled from announce context
(periodic k_timers rescheduling themselves, callbacks starting new
timers, slice-timer rearm during ready-thread wakeup). All other
callers -- k_sleep(), z_abort_timeout(), initial k_timer_start() from
a thread, k_sched_time_slice_set() -- continue to use the +1 round-up.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-17 00:36:30 -04:00
|
|
|
*/
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
dticks = timeout.ticks + ticks_elapsed +
|
|
|
|
|
(this_cpu_announcing() ? 0 : 1);
|
|
|
|
|
ticks = curr_tick + dticks;
|
2025-02-25 16:02:46 -08:00
|
|
|
} else {
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
dticks = Z_TICK_ABS(timeout.ticks) - curr_tick;
|
|
|
|
|
dticks = max(1, dticks);
|
2025-04-04 09:33:03 +02:00
|
|
|
ticks = timeout.ticks;
|
2021-05-24 11:24:13 +02:00
|
|
|
}
|
|
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
if (z_timeout_q_insert(to, dticks) && !any_cpu_announcing()) {
|
2025-03-31 13:25:03 +02:00
|
|
|
if (!has_elapsed) {
|
|
|
|
|
/* In case of absolute timeout that is first to expire
|
|
|
|
|
* elapsed need to be read from the system clock.
|
|
|
|
|
*/
|
|
|
|
|
ticks_elapsed = elapsed();
|
|
|
|
|
}
|
2026-07-22 09:25:43 -04:00
|
|
|
sys_clock_set_timeout(next_timeout(ticks_elapsed), false);
|
2018-11-22 11:49:32 +01:00
|
|
|
}
|
2018-11-20 08:26:34 -08:00
|
|
|
}
|
2025-04-04 09:33:03 +02:00
|
|
|
|
|
|
|
|
return ticks;
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
|
2026-05-27 13:18:00 -04:00
|
|
|
int z_try_abort_timeout(struct _timeout *to)
|
|
|
|
|
{
|
|
|
|
|
int ret = -EINVAL;
|
|
|
|
|
|
|
|
|
|
K_SPINLOCK(&timeout_lock) {
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
if (!z_is_inactive_timeout(to)) {
|
|
|
|
|
bool was_first = z_timeout_q_remove(to);
|
2026-05-27 13:18:00 -04:00
|
|
|
|
|
|
|
|
ret = 0;
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
if (was_first) {
|
2026-07-30 18:01:00 -04:00
|
|
|
reprogram_next(elapsed());
|
2026-05-27 13:18:00 -04:00
|
|
|
}
|
kernel: timeout: tag inflight_timeout with a superseded bit
When z_try_abort_timeout() finds the target timeout already popped
from the queue and in flight (its handler dispatching), the abort
cannot remove anything from the queue. The handler is about to run
(or is running) the timeout's callback; the aborter needs a way to
tell it "you were aborted, skip your side effects". main carried that
signal in the per-timeout dticks=ABORTED sentinel, which any CPU
could write under timeout_lock and the handler checked at entry via
z_is_timeout_handler_canceled(). Later commits in this series remove
the dticks-cancel mechanism, so the signal needs a new home.
Encode it in the low bit of the file-local inflight_timeout pointer
(struct _timeout is pointer-aligned, so bit 0 is free):
inflight_timeout == NULL no handler in flight
inflight_timeout == t handler in flight, not superseded
inflight_timeout == t | 1 handler in flight, superseded
z_try_abort_timeout() sets the bit whenever the target is the
in-flight timeout -- on the announcing CPU (same-CPU IRQ that
preempted the dispatch, or a stop after a re-arm) and on another CPU
racing the handler.
A handler with non-idempotent side effects (currently only k_timer's
expiry_fn) checks z_timeout_inflight_superseded() at entry and bails
if set. Idempotent handlers (z_thread_timeout, work, poll, ...)
tolerate the race and don't need the check.
The bit is a best-effort signal, not a barrier: an aborter that sets
it after the handler has passed its check has no effect (the handler
already committed), exactly as the dticks sentinel behaved on main.
The -EAGAIN return for the cross-CPU case is a separate mechanism,
used only by callers that must wait for the handler to fully complete
(e.g. before freeing the timeout's storage); they spin, while
best-effort callers ignore -EAGAIN and rely on the bit.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
|
|
|
} else if (IS_ENABLED(CONFIG_SMP) && inflight_ptr() == to &&
|
|
|
|
|
!this_cpu_announcing()) {
|
|
|
|
|
/* Handler in flight on another CPU. Free-safety
|
|
|
|
|
* callers retry on -EAGAIN to wait it out; others
|
|
|
|
|
* rely on the superseded mark below and don't.
|
|
|
|
|
*/
|
|
|
|
|
ret = -EAGAIN;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
/* Record that the in-flight timeout has been aborted, so a
|
|
|
|
|
* handler that checks (z_timeout_inflight_superseded) can
|
|
|
|
|
* tell its dispatch was overtaken.
|
|
|
|
|
*/
|
|
|
|
|
if (inflight_ptr() == to) {
|
|
|
|
|
inflight_mark_superseded();
|
2026-05-27 13:18:00 -04:00
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-02 19:33:50 -04:00
|
|
|
if (IS_ENABLED(CONFIG_SMP) && ret == -EAGAIN) {
|
2026-05-27 13:18:00 -04:00
|
|
|
arch_spin_relax();
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
return ret;
|
|
|
|
|
}
|
|
|
|
|
|
kernel: timeout: tag inflight_timeout with a superseded bit
When z_try_abort_timeout() finds the target timeout already popped
from the queue and in flight (its handler dispatching), the abort
cannot remove anything from the queue. The handler is about to run
(or is running) the timeout's callback; the aborter needs a way to
tell it "you were aborted, skip your side effects". main carried that
signal in the per-timeout dticks=ABORTED sentinel, which any CPU
could write under timeout_lock and the handler checked at entry via
z_is_timeout_handler_canceled(). Later commits in this series remove
the dticks-cancel mechanism, so the signal needs a new home.
Encode it in the low bit of the file-local inflight_timeout pointer
(struct _timeout is pointer-aligned, so bit 0 is free):
inflight_timeout == NULL no handler in flight
inflight_timeout == t handler in flight, not superseded
inflight_timeout == t | 1 handler in flight, superseded
z_try_abort_timeout() sets the bit whenever the target is the
in-flight timeout -- on the announcing CPU (same-CPU IRQ that
preempted the dispatch, or a stop after a re-arm) and on another CPU
racing the handler.
A handler with non-idempotent side effects (currently only k_timer's
expiry_fn) checks z_timeout_inflight_superseded() at entry and bails
if set. Idempotent handlers (z_thread_timeout, work, poll, ...)
tolerate the race and don't need the check.
The bit is a best-effort signal, not a barrier: an aborter that sets
it after the handler has passed its check has no effect (the handler
already committed), exactly as the dticks sentinel behaved on main.
The -EAGAIN return for the cross-CPU case is a separate mechanism,
used only by callers that must wait for the handler to fully complete
(e.g. before freeing the timeout's storage); they spin, while
best-effort callers ignore -EAGAIN and rely on the bit.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
|
|
|
bool z_timeout_inflight_superseded(const struct _timeout *to)
|
|
|
|
|
{
|
|
|
|
|
bool superseded = false;
|
|
|
|
|
|
|
|
|
|
K_SPINLOCK(&timeout_lock) {
|
|
|
|
|
superseded = inflight_timeout ==
|
|
|
|
|
(struct _timeout *)((uintptr_t)to | INFLIGHT_SUPERSEDED_BIT);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
return superseded;
|
|
|
|
|
}
|
|
|
|
|
|
2020-09-18 16:24:57 -05:00
|
|
|
k_ticks_t z_timeout_remaining(const struct _timeout *timeout)
|
2020-03-09 13:59:15 -07:00
|
|
|
{
|
|
|
|
|
k_ticks_t ticks = 0;
|
|
|
|
|
|
2023-07-07 09:12:38 +02:00
|
|
|
K_SPINLOCK(&timeout_lock) {
|
2024-03-06 11:12:42 -05:00
|
|
|
if (!z_is_inactive_timeout(timeout)) {
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
ticks = z_timeout_q_remainder(timeout) - elapsed();
|
2024-03-06 11:12:42 -05:00
|
|
|
}
|
2020-03-09 13:59:15 -07:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
return ticks;
|
|
|
|
|
}
|
2025-12-09 10:09:23 -08:00
|
|
|
EXPORT_SYMBOL(z_timeout_remaining);
|
2020-03-09 13:59:15 -07:00
|
|
|
|
2020-09-18 16:24:57 -05:00
|
|
|
k_ticks_t z_timeout_expires(const struct _timeout *timeout)
|
2020-03-09 13:59:15 -07:00
|
|
|
{
|
|
|
|
|
k_ticks_t ticks = 0;
|
|
|
|
|
|
2023-07-07 09:12:38 +02:00
|
|
|
K_SPINLOCK(&timeout_lock) {
|
2024-03-06 11:12:42 -05:00
|
|
|
ticks = curr_tick;
|
|
|
|
|
if (!z_is_inactive_timeout(timeout)) {
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
ticks += z_timeout_q_remainder(timeout);
|
2024-03-06 11:12:42 -05:00
|
|
|
}
|
2020-03-09 13:59:15 -07:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
return ticks;
|
|
|
|
|
}
|
2025-12-09 10:09:23 -08:00
|
|
|
EXPORT_SYMBOL(z_timeout_expires);
|
2020-03-09 13:59:15 -07:00
|
|
|
|
kernel: timeout: report the no-deadline case from the idle path
The idle path asks this for the time until the next wakeup, and gets a
tick count that never says "there is nothing to wake up for": an empty
timeout list is reported as the capped announce budget, exactly like a
deadline further out than can be programmed in one step. The power
management code cannot then tell the two apart, so it arms the timer in
both cases and a system with sloppy idle enabled keeps waking up for
nothing.
Make the same decision reprogram_next() makes. With
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE allowing the uptime to drift, an empty
list is reported as K_TICKS_FOREVER; anything else is a wait, whether a
real deadline or the synthetic one that keeps the announce range
covered. Without sloppy idle the empty case stays a wait, so the timer
remains armed and timekeeping is unaffected.
The return type becomes unsigned to match the tick type used throughout
the timer interface. The conversion at the only caller, in the idle
path, is value preserving in both directions: every wait is capped at
SYS_CLOCK_MAX_WAIT, which is INT32_MAX, and K_TICKS_FOREVER is the same
value read either way. Nothing outside the kernel sees the change; the
power management API keeps its signed tick counts.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-08-01 15:18:07 -04:00
|
|
|
uint32_t z_get_next_timeout_expiry(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
kernel: timeout: report the no-deadline case from the idle path
The idle path asks this for the time until the next wakeup, and gets a
tick count that never says "there is nothing to wake up for": an empty
timeout list is reported as the capped announce budget, exactly like a
deadline further out than can be programmed in one step. The power
management code cannot then tell the two apart, so it arms the timer in
both cases and a system with sloppy idle enabled keeps waking up for
nothing.
Make the same decision reprogram_next() makes. With
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE allowing the uptime to drift, an empty
list is reported as K_TICKS_FOREVER; anything else is a wait, whether a
real deadline or the synthetic one that keeps the announce range
covered. Without sloppy idle the empty case stays a wait, so the timer
remains armed and timekeeping is unaffected.
The return type becomes unsigned to match the tick type used throughout
the timer interface. The conversion at the only caller, in the idle
path, is value preserving in both directions: every wait is capped at
SYS_CLOCK_MAX_WAIT, which is INT32_MAX, and K_TICKS_FOREVER is the same
value read either way. Nothing outside the kernel sees the change; the
power management API keeps its signed tick counts.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-08-01 15:18:07 -04:00
|
|
|
uint32_t ret = (uint32_t)K_TICKS_FOREVER;
|
2018-09-27 16:50:00 -07:00
|
|
|
|
2023-07-07 09:12:38 +02:00
|
|
|
K_SPINLOCK(&timeout_lock) {
|
2026-06-10 17:03:30 -04:00
|
|
|
/*
|
kernel: timeout: report the no-deadline case from the idle path
The idle path asks this for the time until the next wakeup, and gets a
tick count that never says "there is nothing to wake up for": an empty
timeout list is reported as the capped announce budget, exactly like a
deadline further out than can be programmed in one step. The power
management code cannot then tell the two apart, so it arms the timer in
both cases and a system with sloppy idle enabled keeps waking up for
nothing.
Make the same decision reprogram_next() makes. With
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE allowing the uptime to drift, an empty
list is reported as K_TICKS_FOREVER; anything else is a wait, whether a
real deadline or the synthetic one that keeps the announce range
covered. Without sloppy idle the empty case stays a wait, so the timer
remains armed and timekeeping is unaffected.
The return type becomes unsigned to match the tick type used throughout
the timer interface. The conversion at the only caller, in the idle
path, is value preserving in both directions: every wait is capped at
SYS_CLOCK_MAX_WAIT, which is INT32_MAX, and K_TICKS_FOREVER is the same
value read either way. Nothing outside the kernel sees the change; the
power management API keeps its signed tick counts.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-08-01 15:18:07 -04:00
|
|
|
* Same decision as reprogram_next(). Sloppy idle lets the
|
|
|
|
|
* uptime drift, so an empty list means nothing to wake up for
|
|
|
|
|
* and that is reported as K_TICKS_FOREVER. Otherwise the
|
|
|
|
|
* answer is a wait: either a real deadline, or the synthetic
|
|
|
|
|
* one that keeps the announce range covered.
|
2026-06-10 17:03:30 -04:00
|
|
|
*/
|
kernel: timeout: report the no-deadline case from the idle path
The idle path asks this for the time until the next wakeup, and gets a
tick count that never says "there is nothing to wake up for": an empty
timeout list is reported as the capped announce budget, exactly like a
deadline further out than can be programmed in one step. The power
management code cannot then tell the two apart, so it arms the timer in
both cases and a system with sloppy idle enabled keeps waking up for
nothing.
Make the same decision reprogram_next() makes. With
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE allowing the uptime to drift, an empty
list is reported as K_TICKS_FOREVER; anything else is a wait, whether a
real deadline or the synthetic one that keeps the announce range
covered. Without sloppy idle the empty case stays a wait, so the timer
remains armed and timekeeping is unaffected.
The return type becomes unsigned to match the tick type used throughout
the timer interface. The conversion at the only caller, in the idle
path, is value preserving in both directions: every wait is capped at
SYS_CLOCK_MAX_WAIT, which is INT32_MAX, and K_TICKS_FOREVER is the same
value read either way. Nothing outside the kernel sees the change; the
power management API keeps its signed tick counts.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-08-01 15:18:07 -04:00
|
|
|
if (IS_ENABLED(CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE) &&
|
|
|
|
|
z_timeout_q_next_expiry() == K_TICKS_FOREVER) {
|
|
|
|
|
ret = (uint32_t)K_TICKS_FOREVER;
|
|
|
|
|
} else {
|
|
|
|
|
ret = next_timeout(elapsed());
|
2026-06-10 17:03:30 -04:00
|
|
|
}
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
return ret;
|
|
|
|
|
}
|
|
|
|
|
|
2026-06-10 14:42:13 -04:00
|
|
|
void sys_clock_announce_locked(uint32_t ticks, k_spinlock_key_t key)
|
2018-12-20 09:23:31 -08:00
|
|
|
{
|
2022-04-12 09:52:39 -07:00
|
|
|
/* We release the lock around the callbacks below, so on SMP
|
|
|
|
|
* systems someone might be already running the loop. Don't
|
2024-07-06 01:12:07 +07:00
|
|
|
* race (which will cause parallel execution of "sequential"
|
2022-04-12 09:52:39 -07:00
|
|
|
* timeouts and confuse apps), just increment the tick count
|
|
|
|
|
* and return.
|
|
|
|
|
*/
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
if (IS_ENABLED(CONFIG_SMP) && any_cpu_announcing()) {
|
2022-04-12 09:52:39 -07:00
|
|
|
announce_remaining += ticks;
|
|
|
|
|
k_spin_unlock(&timeout_lock, key);
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
|
2018-12-20 09:23:31 -08:00
|
|
|
announce_remaining = ticks;
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
announcing_cpu = CPU_ID;
|
|
|
|
|
|
kernel: timeout: add timer-wheel backend
Add a hierarchical timer wheel as a third selectable timeout backend.
Timeouts are bucketed by expiry distance: one list per tick for the
next 32 ticks ("soon"), one list per 32-tick band for the next ~1024
("later"), and a sorted overflow list ("distant"). Insertion and
removal are O(1) for the common near-future case; every 32 ticks the
announce path sifts the current "later" band into "soon" and refills
it from "distant". This scales well when many short-lived timeouts are
pending.
Unlike the dlist and min-heap backends, the wheel does not fit the
generic next_gap/advance/pop_due announce primitives: its per-tick
advance is a bitmap scan that jumps over empty ticks, and its sift is
a time-driven event tied to no single timeout. Rather than contort
that (already subtle) state machine, the wheel uses the backend-owned
announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and
implements z_timeout_q_announce(), which sys_clock_announce_locked()
calls in place of the generic loop.
Like the other backends the wheel is a single implementation header
(kernel/timeout_wheel.h) included only by timeout.c, so its state, its
operations and that announce loop reach the shared state (curr_tick,
announce_remaining, inflight_timeout) directly; nothing extra needs
exposing. The SMP re-entry guard, announcing_cpu and the reprogram
remain in timeout.c. The wheel fires handlers through the same
inflight_timeout dance, so the post-#109977 abort/in-flight
synchronization works unchanged, and it carries no per-node
ANNOUNCING/ABORTED sentinels.
struct _timeout grows a wheel-only flags field (which wheel tier a
timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist
and heap builds are unaffected.
The backend is EXPERIMENTAL. Two known limitations, inherent to the
wheel algorithm (not the abstraction):
- No same-tick firing-order guarantee (sifted timeouts are
prepended), so the timeout_order test does not apply.
- next_timeout() never exceeds 32 ticks because a sift is always
pending, so the wheel wakes a tickless-idle CPU at least every 32
ticks. This fails tests that assert zero spurious idle wakeups
(tests/kernel/context cpu_idle / timer_interrupts) and is a power
regression versus the dlist and heap backends.
Verified the timeout-functional suites (timer_api, timeout,
timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86,
x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected.
The timer-wheel data structure, bucketing scheme and sift algorithm
are the work of Peter Mitsis (PR #108339), re-homed here behind the
timeout backend interface.
Co-authored-by: Peter Mitsis <peter.mitsis@intel.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
|
|
|
#ifdef _TIMEOUT_BACKEND_OWNS_ANNOUNCE
|
|
|
|
|
/* The backend (timer wheel) runs the firing loop itself: its per-tick
|
|
|
|
|
* advance is a bitmap jump over empty ticks and its sift is a
|
|
|
|
|
* time-driven event with no single timeout to drive a generic loop.
|
|
|
|
|
* It advances curr_tick / announce_remaining and fires handlers via
|
|
|
|
|
* the same inflight_timeout dance as below.
|
|
|
|
|
*/
|
|
|
|
|
key = z_timeout_q_announce(key);
|
|
|
|
|
#else
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
while (announce_remaining > 0) {
|
|
|
|
|
struct _timeout *t;
|
|
|
|
|
int32_t dt = z_timeout_q_next_gap();
|
|
|
|
|
|
|
|
|
|
if ((uint32_t)dt > announce_remaining) {
|
|
|
|
|
/* Next event is past this announce window: advance the
|
|
|
|
|
* backend by the residual ticks without firing anything.
|
|
|
|
|
*/
|
|
|
|
|
dt = (int32_t)announce_remaining;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
/* Advance curr_tick and decrement announce_remaining together
|
|
|
|
|
* under the lock so non-announcing CPUs observe a consistent
|
|
|
|
|
* (curr_tick + announce_remaining + sys_clock_elapsed()) ==
|
|
|
|
|
* T_real even while we drop the lock around handlers. The "we
|
|
|
|
|
* are announcing" state is carried by announcing_cpu, so
|
|
|
|
|
* announce_remaining reaching 0 mid-loop is harmless: same-tick
|
|
|
|
|
* handlers' z_add_timeout() still anchors via
|
|
|
|
|
* this_cpu_announcing(), and another CPU's announce still folds
|
|
|
|
|
* into ours via any_cpu_announcing() in the SMP early-return.
|
kernel: timeout: keep announce_remaining stable across same-tick group
When sys_clock_announce_locked() processes a tick that has multiple
timeouts queued for it, the timeout queue stores the second and
subsequent ones with dticks == 0 (relative to the first). The original
loop fired each in turn and decremented announce_remaining at the
bottom of every iteration:
announce_remaining -= dt;
For the first timeout in a same-tick group, dt is the cumulative tick
delta that brought us to that tick. After that subtraction
announce_remaining can drop to zero, even though there are still
same-tick callbacks queued for the loop to fire. Each subsequent
same-tick callback then runs while announce_remaining == 0, which
breaks two invariants the rest of the kernel relies on:
* The SMP early-return at the top of sys_clock_announce_locked():
if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) {
announce_remaining += ticks;
k_spin_unlock(&timeout_lock, key);
return;
}
is meant to detect that another CPU is already inside the loop and
fold the new ticks into the ongoing announce. The lock is released
around each callback, so during a same-tick callback another CPU
can grab the lock, see announce_remaining == 0, miss the early
return, set announce_remaining = ticks of its own, and start
walking the queue in parallel with the original announcer -- the
exact race the early return is supposed to prevent.
* The elapsed() helper:
return announce_remaining == 0 ? sys_clock_elapsed() : 0U;
is meant to return 0 for any z_add_timeout() / z_abort_timeout()
call that happens from inside a tick-processing callback, so that
timeouts scheduled from such a callback are anchored to the
currently-firing tick rather than to a fresh sys_clock_elapsed()
reading. With announce_remaining == 0 mid-loop, two callbacks on
the same tick observe inconsistent semantics: the first one (that
saw announce_remaining > 0) gets dticks anchored to the firing
tick, while subsequent ones (seeing 0) get dticks computed against
a fresh elapsed() reading and end up off by one tick. Two periodic
timers that happen to fire on the same tick will therefore
permanently drift apart by one tick going forward.
Restructure the loop so the announce_remaining decrement happens once
per tick rather than once per timeout: the outer while drives forward
across distinct ticks, and an inner do-while drains all timeouts
queued on the current tick before announce_remaining is updated.
announce_remaining now stays at its pre-tick value for the entire
same-tick group, which both the SMP early return and elapsed()
correctly observe as non-zero.
remove_timeout()'s dticks propagation is also unnecessary in this
loop because curr_tick is advanced by t->dticks before the timeout
is unlinked, which keeps the next item's stored dticks valid relative
to the new curr_tick. sys_dlist_remove() on its own is sufficient.
Loop structure suggested by Peter Mitsis.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
|
|
|
*/
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
z_timeout_q_advance(dt);
|
|
|
|
|
curr_tick += dt;
|
|
|
|
|
announce_remaining -= dt;
|
kernel: timeout: keep announce_remaining stable across same-tick group
When sys_clock_announce_locked() processes a tick that has multiple
timeouts queued for it, the timeout queue stores the second and
subsequent ones with dticks == 0 (relative to the first). The original
loop fired each in turn and decremented announce_remaining at the
bottom of every iteration:
announce_remaining -= dt;
For the first timeout in a same-tick group, dt is the cumulative tick
delta that brought us to that tick. After that subtraction
announce_remaining can drop to zero, even though there are still
same-tick callbacks queued for the loop to fire. Each subsequent
same-tick callback then runs while announce_remaining == 0, which
breaks two invariants the rest of the kernel relies on:
* The SMP early-return at the top of sys_clock_announce_locked():
if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) {
announce_remaining += ticks;
k_spin_unlock(&timeout_lock, key);
return;
}
is meant to detect that another CPU is already inside the loop and
fold the new ticks into the ongoing announce. The lock is released
around each callback, so during a same-tick callback another CPU
can grab the lock, see announce_remaining == 0, miss the early
return, set announce_remaining = ticks of its own, and start
walking the queue in parallel with the original announcer -- the
exact race the early return is supposed to prevent.
* The elapsed() helper:
return announce_remaining == 0 ? sys_clock_elapsed() : 0U;
is meant to return 0 for any z_add_timeout() / z_abort_timeout()
call that happens from inside a tick-processing callback, so that
timeouts scheduled from such a callback are anchored to the
currently-firing tick rather than to a fresh sys_clock_elapsed()
reading. With announce_remaining == 0 mid-loop, two callbacks on
the same tick observe inconsistent semantics: the first one (that
saw announce_remaining > 0) gets dticks anchored to the firing
tick, while subsequent ones (seeing 0) get dticks computed against
a fresh elapsed() reading and end up off by one tick. Two periodic
timers that happen to fire on the same tick will therefore
permanently drift apart by one tick going forward.
Restructure the loop so the announce_remaining decrement happens once
per tick rather than once per timeout: the outer while drives forward
across distinct ticks, and an inner do-while drains all timeouts
queued on the current tick before announce_remaining is updated.
announce_remaining now stays at its pre-tick value for the entire
same-tick group, which both the SMP early return and elapsed()
correctly observe as non-zero.
remove_timeout()'s dticks propagation is also unnecessary in this
loop because curr_tick is advanced by t->dticks before the timeout
is unlinked, which keeps the next item's stored dticks valid relative
to the new curr_tick. sys_dlist_remove() on its own is sufficient.
Loop structure suggested by Peter Mitsis.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
/* Drain everything due at the new curr_tick. */
|
|
|
|
|
while ((t = z_timeout_q_pop_due()) != NULL) {
|
|
|
|
|
_timeout_func_t handler = t->fn;
|
kernel: timeout: keep announce_remaining stable across same-tick group
When sys_clock_announce_locked() processes a tick that has multiple
timeouts queued for it, the timeout queue stores the second and
subsequent ones with dticks == 0 (relative to the first). The original
loop fired each in turn and decremented announce_remaining at the
bottom of every iteration:
announce_remaining -= dt;
For the first timeout in a same-tick group, dt is the cumulative tick
delta that brought us to that tick. After that subtraction
announce_remaining can drop to zero, even though there are still
same-tick callbacks queued for the loop to fire. Each subsequent
same-tick callback then runs while announce_remaining == 0, which
breaks two invariants the rest of the kernel relies on:
* The SMP early-return at the top of sys_clock_announce_locked():
if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) {
announce_remaining += ticks;
k_spin_unlock(&timeout_lock, key);
return;
}
is meant to detect that another CPU is already inside the loop and
fold the new ticks into the ongoing announce. The lock is released
around each callback, so during a same-tick callback another CPU
can grab the lock, see announce_remaining == 0, miss the early
return, set announce_remaining = ticks of its own, and start
walking the queue in parallel with the original announcer -- the
exact race the early return is supposed to prevent.
* The elapsed() helper:
return announce_remaining == 0 ? sys_clock_elapsed() : 0U;
is meant to return 0 for any z_add_timeout() / z_abort_timeout()
call that happens from inside a tick-processing callback, so that
timeouts scheduled from such a callback are anchored to the
currently-firing tick rather than to a fresh sys_clock_elapsed()
reading. With announce_remaining == 0 mid-loop, two callbacks on
the same tick observe inconsistent semantics: the first one (that
saw announce_remaining > 0) gets dticks anchored to the firing
tick, while subsequent ones (seeing 0) get dticks computed against
a fresh elapsed() reading and end up off by one tick. Two periodic
timers that happen to fire on the same tick will therefore
permanently drift apart by one tick going forward.
Restructure the loop so the announce_remaining decrement happens once
per tick rather than once per timeout: the outer while drives forward
across distinct ticks, and an inner do-while drains all timeouts
queued on the current tick before announce_remaining is updated.
announce_remaining now stays at its pre-tick value for the entire
same-tick group, which both the SMP early return and elapsed()
correctly observe as non-zero.
remove_timeout()'s dticks propagation is also unnecessary in this
loop because curr_tick is advanced by t->dticks before the timeout
is unlinked, which keeps the next item's stored dticks valid relative
to the new curr_tick. sys_dlist_remove() on its own is sufficient.
Loop structure suggested by Peter Mitsis.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
inflight_timeout = t;
|
2018-12-20 09:23:31 -08:00
|
|
|
|
kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.
Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:
- The per-node helpers z_init_timeout() / z_is_inactive_timeout()
operate on a single struct _timeout and are needed tree-wide
(timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
depend only on struct _timeout's fields.
- The queue itself (its instance and the z_timeout_q_*() operations)
is used only by timeout.c, so it becomes a backend implementation
header (kernel/timeout_list.h) that timeout.c includes directly,
after its shared state. The header is private to that translation
unit; no other file sees the queue or its instance, so the instance
is plain static and needs no extern.
timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.
The backend exposes:
- z_timeout_q_insert / _remove add and arbitrary remove
- z_timeout_q_remainder / _next_timeout queries
- z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives
The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.
So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.
The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.
Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
|
|
|
k_spin_unlock(&timeout_lock, key);
|
|
|
|
|
handler(t);
|
|
|
|
|
key = k_spin_lock(&timeout_lock);
|
|
|
|
|
inflight_timeout = NULL;
|
|
|
|
|
}
|
2018-12-20 09:23:31 -08:00
|
|
|
}
|
kernel: timeout: add timer-wheel backend
Add a hierarchical timer wheel as a third selectable timeout backend.
Timeouts are bucketed by expiry distance: one list per tick for the
next 32 ticks ("soon"), one list per 32-tick band for the next ~1024
("later"), and a sorted overflow list ("distant"). Insertion and
removal are O(1) for the common near-future case; every 32 ticks the
announce path sifts the current "later" band into "soon" and refills
it from "distant". This scales well when many short-lived timeouts are
pending.
Unlike the dlist and min-heap backends, the wheel does not fit the
generic next_gap/advance/pop_due announce primitives: its per-tick
advance is a bitmap scan that jumps over empty ticks, and its sift is
a time-driven event tied to no single timeout. Rather than contort
that (already subtle) state machine, the wheel uses the backend-owned
announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and
implements z_timeout_q_announce(), which sys_clock_announce_locked()
calls in place of the generic loop.
Like the other backends the wheel is a single implementation header
(kernel/timeout_wheel.h) included only by timeout.c, so its state, its
operations and that announce loop reach the shared state (curr_tick,
announce_remaining, inflight_timeout) directly; nothing extra needs
exposing. The SMP re-entry guard, announcing_cpu and the reprogram
remain in timeout.c. The wheel fires handlers through the same
inflight_timeout dance, so the post-#109977 abort/in-flight
synchronization works unchanged, and it carries no per-node
ANNOUNCING/ABORTED sentinels.
struct _timeout grows a wheel-only flags field (which wheel tier a
timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist
and heap builds are unaffected.
The backend is EXPERIMENTAL. Two known limitations, inherent to the
wheel algorithm (not the abstraction):
- No same-tick firing-order guarantee (sifted timeouts are
prepended), so the timeout_order test does not apply.
- next_timeout() never exceeds 32 ticks because a sift is always
pending, so the wheel wakes a tickless-idle CPU at least every 32
ticks. This fails tests that assert zero spurious idle wakeups
(tests/kernel/context cpu_idle / timer_interrupts) and is a power
regression versus the dlist and heap backends.
Verified the timeout-functional suites (timer_api, timeout,
timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86,
x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected.
The timer-wheel data structure, bucketing scheme and sift algorithm
are the work of Peter Mitsis (PR #108339), re-homed here behind the
timeout backend interface.
Co-authored-by: Peter Mitsis <peter.mitsis@intel.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
|
|
|
#endif /* _TIMEOUT_BACKEND_OWNS_ANNOUNCE */
|
2018-12-20 09:23:31 -08:00
|
|
|
|
kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.
Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.
Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).
elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is
curr_tick + announce_remaining + sys_clock_elapsed() == T_real
on every CPU at every point inside or outside the announce window.
Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.
Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:
curr_tick += t->dticks;
announce_remaining -= t->dticks;
Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.
Fixes #106317
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
|
|
|
announcing_cpu = -1;
|
2018-12-20 09:23:31 -08:00
|
|
|
|
2026-07-30 18:01:00 -04:00
|
|
|
reprogram_next(0);
|
2018-12-20 09:23:31 -08:00
|
|
|
|
|
|
|
|
k_spin_unlock(&timeout_lock, key);
|
2023-03-06 14:31:35 -08:00
|
|
|
|
|
|
|
|
#ifdef CONFIG_TIMESLICING
|
|
|
|
|
z_time_slice();
|
2024-03-08 12:00:10 +01:00
|
|
|
#endif /* CONFIG_TIMESLICING */
|
2018-12-20 09:23:31 -08:00
|
|
|
}
|
|
|
|
|
|
kernel/timeout: introduce sys_clock_lock() and sys_clock_announce_locked()
On SMP systems with tickless kernels, a race condition exists between
timer driver ISRs and the kernel's tick accounting. The driver updates
its hardware cycle baseline under a private lock, then calls
sys_clock_announce() which updates curr_tick under the separate
timeout_lock. In the gap between these two lock releases, any kernel
code calling sys_clock_elapsed() sees the new driver baseline but the
old curr_tick, producing inconsistent time values that can go backwards.
This affects every code path using the internal elapsed() helper:
uptime queries, timeout scheduling, timeout cancellation, remaining
time queries, and next-expiry calculations.
The root cause is two separate locks protecting state that must be
mutually consistent. Fix this by exposing the kernel's timeout_lock
to timer drivers via sys_clock_lock()/sys_clock_unlock(), and
providing sys_clock_announce_locked() which assumes the lock is
already held.
Timer drivers can now acquire the single lock, update their hardware
state, and announce ticks all under the same lock — eliminating the
race window entirely. The key is passed to sys_clock_announce_locked()
which consumes it (releasing the lock when it returns).
The existing sys_clock_announce() becomes a backward-compatible wrapper,
allowing incremental driver migration with no flag day.
Document that sys_clock_set_timeout(), sys_clock_elapsed(), and
sys_clock_idle_exit() are called by the kernel with the timer lock
held. Update the timer driver guide in clocks.rst accordingly.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-03-13 16:43:25 -04:00
|
|
|
#if defined(CONFIG_SMP) || defined(CONFIG_SPIN_VALIDATE)
|
|
|
|
|
k_spinlock_key_t sys_clock_lock(void)
|
|
|
|
|
{
|
|
|
|
|
return k_spin_lock(&timeout_lock);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void sys_clock_unlock(k_spinlock_key_t key)
|
|
|
|
|
{
|
|
|
|
|
k_spin_unlock(&timeout_lock, key);
|
|
|
|
|
}
|
|
|
|
|
#endif
|
|
|
|
|
|
2026-04-19 00:42:18 -04:00
|
|
|
#if defined(CONFIG_TEST) || defined(CONFIG_ASSERT)
|
|
|
|
|
bool sys_clock_is_locked(void)
|
|
|
|
|
{
|
|
|
|
|
return z_spin_is_locked(&timeout_lock);
|
|
|
|
|
}
|
|
|
|
|
#endif
|
|
|
|
|
|
2021-03-13 08:21:21 -05:00
|
|
|
int64_t sys_clock_tick_get(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
2020-05-27 11:26:57 -05:00
|
|
|
uint64_t t = 0U;
|
2018-09-27 16:50:00 -07:00
|
|
|
|
2023-07-07 09:12:38 +02:00
|
|
|
K_SPINLOCK(&timeout_lock) {
|
2022-08-03 16:11:32 -04:00
|
|
|
t = curr_tick + elapsed();
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
return t;
|
|
|
|
|
}
|
|
|
|
|
|
2021-03-13 08:19:53 -05:00
|
|
|
uint32_t sys_clock_tick_get_32(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
2018-10-02 11:12:08 -07:00
|
|
|
#ifdef CONFIG_TICKLESS_KERNEL
|
2021-03-13 08:21:21 -05:00
|
|
|
return (uint32_t)sys_clock_tick_get();
|
2018-10-02 11:12:08 -07:00
|
|
|
#else
|
2020-05-27 11:26:57 -05:00
|
|
|
return (uint32_t)curr_tick;
|
2024-03-08 12:00:10 +01:00
|
|
|
#endif /* CONFIG_TICKLESS_KERNEL */
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
|
2020-05-27 11:26:57 -05:00
|
|
|
int64_t z_impl_k_uptime_ticks(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
2021-03-13 08:21:21 -05:00
|
|
|
return sys_clock_tick_get();
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
#ifdef CONFIG_USERSPACE
|
2020-05-27 11:26:57 -05:00
|
|
|
static inline int64_t z_vrfy_k_uptime_ticks(void)
|
2018-09-27 16:50:00 -07:00
|
|
|
{
|
2020-03-10 15:26:38 -07:00
|
|
|
return z_impl_k_uptime_ticks();
|
2018-09-27 16:50:00 -07:00
|
|
|
}
|
2024-01-24 17:35:04 +08:00
|
|
|
#include <zephyr/syscalls/k_uptime_ticks_mrsh.c>
|
2024-03-08 12:00:10 +01:00
|
|
|
#endif /* CONFIG_USERSPACE */
|
kernel/timeout: Make timeout arguments an opaque type
Add a k_timeout_t type, and use it everywhere that kernel API
functions were accepting a millisecond timeout argument. Instead of
forcing milliseconds everywhere (which are often not integrally
representable as system ticks), do the conversion to ticks at the
point where the timeout is created. This avoids an extra unit
conversion in some application code, and allows us to express the
timeout in units other than milliseconds to achieve greater precision.
The existing K_MSEC() et. al. macros now return initializers for a
k_timeout_t.
The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t
values, which means they cannot be operated on as integers.
Applications which have their own APIs that need to inspect these
vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to
test for equality.
Timer drivers, which receive an integer tick count in ther
z_clock_set_timeout() functions, now use the integer-valued
K_TICKS_FOREVER constant instead of K_FOREVER.
For the initial release, to preserve source compatibility, a
CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the
k_timeout_t will remain a compatible 32 bit value that will work with
any legacy Zephyr application.
Some subsystems present timeout (or timeout-like) values to their own
users as APIs that would re-use the kernel's own constants and
conventions. These will require some minor design work to adapt to
the new scheme (in most cases just using k_timeout_t directly in their
own API), and they have not been changed in this patch, instead
selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems
include: CAN Bus, the Microbit display driver, I2S, LoRa modem
drivers, the UART Async API, Video hardware drivers, the console
subsystem, and the network buffer abstraction.
k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant
provided that works identically to the original API.
Most of the changes here are just type/configuration management and
documentation, but there are logic changes in mempool, where a loop
that used a timeout numerically has been reworked using a new
z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was
enabled) a similar loop was needlessly used to try to retry the
k_poll() call after a spurious failure. But k_poll() does not fail
spuriously, so the loop was removed.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
|
|
|
|
2023-07-06 13:56:01 -04:00
|
|
|
k_timepoint_t sys_timepoint_calc(k_timeout_t timeout)
|
kernel/timeout: Make timeout arguments an opaque type
Add a k_timeout_t type, and use it everywhere that kernel API
functions were accepting a millisecond timeout argument. Instead of
forcing milliseconds everywhere (which are often not integrally
representable as system ticks), do the conversion to ticks at the
point where the timeout is created. This avoids an extra unit
conversion in some application code, and allows us to express the
timeout in units other than milliseconds to achieve greater precision.
The existing K_MSEC() et. al. macros now return initializers for a
k_timeout_t.
The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t
values, which means they cannot be operated on as integers.
Applications which have their own APIs that need to inspect these
vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to
test for equality.
Timer drivers, which receive an integer tick count in ther
z_clock_set_timeout() functions, now use the integer-valued
K_TICKS_FOREVER constant instead of K_FOREVER.
For the initial release, to preserve source compatibility, a
CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the
k_timeout_t will remain a compatible 32 bit value that will work with
any legacy Zephyr application.
Some subsystems present timeout (or timeout-like) values to their own
users as APIs that would re-use the kernel's own constants and
conventions. These will require some minor design work to adapt to
the new scheme (in most cases just using k_timeout_t directly in their
own API), and they have not been changed in this patch, instead
selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems
include: CAN Bus, the Microbit display driver, I2S, LoRa modem
drivers, the UART Async API, Video hardware drivers, the console
subsystem, and the network buffer abstraction.
k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant
provided that works identically to the original API.
Most of the changes here are just type/configuration management and
documentation, but there are logic changes in mempool, where a loop
that used a timeout numerically has been reworked using a new
z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was
enabled) a similar loop was needlessly used to try to retry the
k_poll() call after a spurious failure. But k_poll() does not fail
spuriously, so the loop was removed.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
|
|
|
{
|
2023-07-06 13:56:01 -04:00
|
|
|
k_timepoint_t timepoint;
|
kernel/timeout: Make timeout arguments an opaque type
Add a k_timeout_t type, and use it everywhere that kernel API
functions were accepting a millisecond timeout argument. Instead of
forcing milliseconds everywhere (which are often not integrally
representable as system ticks), do the conversion to ticks at the
point where the timeout is created. This avoids an extra unit
conversion in some application code, and allows us to express the
timeout in units other than milliseconds to achieve greater precision.
The existing K_MSEC() et. al. macros now return initializers for a
k_timeout_t.
The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t
values, which means they cannot be operated on as integers.
Applications which have their own APIs that need to inspect these
vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to
test for equality.
Timer drivers, which receive an integer tick count in ther
z_clock_set_timeout() functions, now use the integer-valued
K_TICKS_FOREVER constant instead of K_FOREVER.
For the initial release, to preserve source compatibility, a
CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the
k_timeout_t will remain a compatible 32 bit value that will work with
any legacy Zephyr application.
Some subsystems present timeout (or timeout-like) values to their own
users as APIs that would re-use the kernel's own constants and
conventions. These will require some minor design work to adapt to
the new scheme (in most cases just using k_timeout_t directly in their
own API), and they have not been changed in this patch, instead
selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems
include: CAN Bus, the Microbit display driver, I2S, LoRa modem
drivers, the UART Async API, Video hardware drivers, the console
subsystem, and the network buffer abstraction.
k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant
provided that works identically to the original API.
Most of the changes here are just type/configuration management and
documentation, but there are logic changes in mempool, where a loop
that used a timeout numerically has been reworked using a new
z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was
enabled) a similar loop was needlessly used to try to retry the
k_poll() call after a spurious failure. But k_poll() does not fail
spuriously, so the loop was removed.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
|
|
|
|
|
|
|
|
if (K_TIMEOUT_EQ(timeout, K_FOREVER)) {
|
2023-07-06 13:56:01 -04:00
|
|
|
timepoint.tick = UINT64_MAX;
|
kernel/timeout: Make timeout arguments an opaque type
Add a k_timeout_t type, and use it everywhere that kernel API
functions were accepting a millisecond timeout argument. Instead of
forcing milliseconds everywhere (which are often not integrally
representable as system ticks), do the conversion to ticks at the
point where the timeout is created. This avoids an extra unit
conversion in some application code, and allows us to express the
timeout in units other than milliseconds to achieve greater precision.
The existing K_MSEC() et. al. macros now return initializers for a
k_timeout_t.
The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t
values, which means they cannot be operated on as integers.
Applications which have their own APIs that need to inspect these
vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to
test for equality.
Timer drivers, which receive an integer tick count in ther
z_clock_set_timeout() functions, now use the integer-valued
K_TICKS_FOREVER constant instead of K_FOREVER.
For the initial release, to preserve source compatibility, a
CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the
k_timeout_t will remain a compatible 32 bit value that will work with
any legacy Zephyr application.
Some subsystems present timeout (or timeout-like) values to their own
users as APIs that would re-use the kernel's own constants and
conventions. These will require some minor design work to adapt to
the new scheme (in most cases just using k_timeout_t directly in their
own API), and they have not been changed in this patch, instead
selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems
include: CAN Bus, the Microbit display driver, I2S, LoRa modem
drivers, the UART Async API, Video hardware drivers, the console
subsystem, and the network buffer abstraction.
k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant
provided that works identically to the original API.
Most of the changes here are just type/configuration management and
documentation, but there are logic changes in mempool, where a loop
that used a timeout numerically has been reworked using a new
z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was
enabled) a similar loop was needlessly used to try to retry the
k_poll() call after a spurious failure. But k_poll() does not fail
spuriously, so the loop was removed.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
|
|
|
} else if (K_TIMEOUT_EQ(timeout, K_NO_WAIT)) {
|
2023-07-06 13:56:01 -04:00
|
|
|
timepoint.tick = 0;
|
2021-03-20 00:36:55 +02:00
|
|
|
} else {
|
2023-07-06 13:56:01 -04:00
|
|
|
k_ticks_t dt = timeout.ticks;
|
2020-03-09 09:35:35 -07:00
|
|
|
|
2025-02-25 16:02:46 -08:00
|
|
|
if (Z_IS_TIMEOUT_RELATIVE(timeout)) {
|
2025-10-21 16:50:55 +01:00
|
|
|
timepoint.tick = sys_clock_tick_get() + max(1, dt);
|
2025-02-25 16:02:46 -08:00
|
|
|
} else {
|
|
|
|
|
timepoint.tick = Z_TICK_ABS(dt);
|
2021-03-20 00:36:55 +02:00
|
|
|
}
|
2020-03-09 09:35:35 -07:00
|
|
|
}
|
2023-07-06 13:56:01 -04:00
|
|
|
|
|
|
|
|
return timepoint;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
k_timeout_t sys_timepoint_timeout(k_timepoint_t timepoint)
|
|
|
|
|
{
|
|
|
|
|
uint64_t now, remaining;
|
|
|
|
|
|
|
|
|
|
if (timepoint.tick == UINT64_MAX) {
|
|
|
|
|
return K_FOREVER;
|
|
|
|
|
}
|
|
|
|
|
if (timepoint.tick == 0) {
|
|
|
|
|
return K_NO_WAIT;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
now = sys_clock_tick_get();
|
|
|
|
|
remaining = (timepoint.tick > now) ? (timepoint.tick - now) : 0;
|
|
|
|
|
return K_TICKS(remaining);
|
kernel/timeout: Make timeout arguments an opaque type
Add a k_timeout_t type, and use it everywhere that kernel API
functions were accepting a millisecond timeout argument. Instead of
forcing milliseconds everywhere (which are often not integrally
representable as system ticks), do the conversion to ticks at the
point where the timeout is created. This avoids an extra unit
conversion in some application code, and allows us to express the
timeout in units other than milliseconds to achieve greater precision.
The existing K_MSEC() et. al. macros now return initializers for a
k_timeout_t.
The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t
values, which means they cannot be operated on as integers.
Applications which have their own APIs that need to inspect these
vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to
test for equality.
Timer drivers, which receive an integer tick count in ther
z_clock_set_timeout() functions, now use the integer-valued
K_TICKS_FOREVER constant instead of K_FOREVER.
For the initial release, to preserve source compatibility, a
CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the
k_timeout_t will remain a compatible 32 bit value that will work with
any legacy Zephyr application.
Some subsystems present timeout (or timeout-like) values to their own
users as APIs that would re-use the kernel's own constants and
conventions. These will require some minor design work to adapt to
the new scheme (in most cases just using k_timeout_t directly in their
own API), and they have not been changed in this patch, instead
selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems
include: CAN Bus, the Microbit display driver, I2S, LoRa modem
drivers, the UART Async API, Video hardware drivers, the console
subsystem, and the network buffer abstraction.
k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant
provided that works identically to the original API.
Most of the changes here are just type/configuration management and
documentation, but there are logic changes in mempool, where a loop
that used a timeout numerically has been reworked using a new
z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was
enabled) a similar loop was needlessly used to try to retry the
k_poll() call after a spurious failure. But k_poll() does not fail
spuriously, so the loop was removed.
Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
|
|
|
}
|
2022-12-15 10:14:43 -08:00
|
|
|
|
|
|
|
|
#ifdef CONFIG_ZTEST
|
|
|
|
|
void z_impl_sys_clock_tick_set(uint64_t tick)
|
|
|
|
|
{
|
|
|
|
|
curr_tick = tick;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void z_vrfy_sys_clock_tick_set(uint64_t tick)
|
|
|
|
|
{
|
|
|
|
|
z_impl_sys_clock_tick_set(tick);
|
|
|
|
|
}
|
2024-03-08 12:00:10 +01:00
|
|
|
#endif /* CONFIG_ZTEST */
|