zephyr/kernel/timeout.c

513 lines
14 KiB
C
Raw Permalink Normal View History

/*
* Copyright (c) 2018 Intel Corporation
*
* SPDX-License-Identifier: Apache-2.0
*/
headers: Refactor kernel and arch headers. This commit refactors kernel and arch headers to establish a boundary between private and public interface headers. The refactoring strategy used in this commit is detailed in the issue This commit introduces the following major changes: 1. Establish a clear boundary between private and public headers by removing "kernel/include" and "arch/*/include" from the global include paths. Ideally, only kernel/ and arch/*/ source files should reference the headers in these directories. If these headers must be used by a component, these include paths shall be manually added to the CMakeLists.txt file of the component. This is intended to discourage applications from including private kernel and arch headers either knowingly and unknowingly. - kernel/include/ (PRIVATE) This directory contains the private headers that provide private kernel definitions which should not be visible outside the kernel and arch source code. All public kernel definitions must be added to an appropriate header located under include/. - arch/*/include/ (PRIVATE) This directory contains the private headers that provide private architecture-specific definitions which should not be visible outside the arch and kernel source code. All public architecture- specific definitions must be added to an appropriate header located under include/arch/*/. - include/ AND include/sys/ (PUBLIC) This directory contains the public headers that provide public kernel definitions which can be referenced by both kernel and application code. - include/arch/*/ (PUBLIC) This directory contains the public headers that provide public architecture-specific definitions which can be referenced by both kernel and application code. 2. Split arch_interface.h into "kernel-to-arch interface" and "public arch interface" divisions. - kernel/include/kernel_arch_interface.h * provides private "kernel-to-arch interface" definition. * includes arch/*/include/kernel_arch_func.h to ensure that the interface function implementations are always available. * includes sys/arch_interface.h so that public arch interface definitions are automatically included when including this file. - arch/*/include/kernel_arch_func.h * provides architecture-specific "kernel-to-arch interface" implementation. * only the functions that will be used in kernel and arch source files are defined here. - include/sys/arch_interface.h * provides "public arch interface" definition. * includes include/arch/arch_inlines.h to ensure that the architecture-specific public inline interface function implementations are always available. - include/arch/arch_inlines.h * includes architecture-specific arch_inlines.h in include/arch/*/arch_inline.h. - include/arch/*/arch_inline.h * provides architecture-specific "public arch interface" inline function implementation. * supersedes include/sys/arch_inline.h. 3. Refactor kernel and the existing architecture implementations. - Remove circular dependency of kernel and arch headers. The following general rules should be observed: * Never include any private headers from public headers * Never include kernel_internal.h in kernel_arch_data.h * Always include kernel_arch_data.h from kernel_arch_func.h * Never include kernel.h from kernel_struct.h either directly or indirectly. Only add the kernel structures that must be referenced from public arch headers in this file. - Relocate syscall_handler.h to include/ so it can be used in the public code. This is necessary because many user-mode public codes reference the functions defined in this header. - Relocate kernel_arch_thread.h to include/arch/*/thread.h. This is necessary to provide architecture-specific thread definition for 'struct k_thread' in kernel.h. - Remove any private header dependencies from public headers using the following methods: * If dependency is not required, simply omit * If dependency is required, - Relocate a portion of the required dependencies from the private header to an appropriate public header OR - Relocate the required private header to make it public. This commit supersedes #20047, addresses #19666, and fixes #3056. Signed-off-by: Stephanos Ioannidis <root@stephanos.io>
2019-10-25 00:08:21 +09:00
#include <zephyr/sys/minmax.h>
#include <zephyr/kernel.h>
#include <zephyr/spinlock.h>
#include <ksched.h>
#include <timeout_q.h>
#include <zephyr/internal/syscall_handler.h>
#include <zephyr/drivers/timer/system_timer.h>
#include <zephyr/sys/clock.h>
#include <zephyr/llext/symbol.h>
#include <timeslicing.h>
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
/* Absolute tick counter: ticks since boot. */
static uint64_t curr_tick;
/*
* The timeout code shall take no locks other than its own (timeout_lock), nor
kernel: timeout: add timer-wheel backend Add a hierarchical timer wheel as a third selectable timeout backend. Timeouts are bucketed by expiry distance: one list per tick for the next 32 ticks ("soon"), one list per 32-tick band for the next ~1024 ("later"), and a sorted overflow list ("distant"). Insertion and removal are O(1) for the common near-future case; every 32 ticks the announce path sifts the current "later" band into "soon" and refills it from "distant". This scales well when many short-lived timeouts are pending. Unlike the dlist and min-heap backends, the wheel does not fit the generic next_gap/advance/pop_due announce primitives: its per-tick advance is a bitmap scan that jumps over empty ticks, and its sift is a time-driven event tied to no single timeout. Rather than contort that (already subtle) state machine, the wheel uses the backend-owned announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and implements z_timeout_q_announce(), which sys_clock_announce_locked() calls in place of the generic loop. Like the other backends the wheel is a single implementation header (kernel/timeout_wheel.h) included only by timeout.c, so its state, its operations and that announce loop reach the shared state (curr_tick, announce_remaining, inflight_timeout) directly; nothing extra needs exposing. The SMP re-entry guard, announcing_cpu and the reprogram remain in timeout.c. The wheel fires handlers through the same inflight_timeout dance, so the post-#109977 abort/in-flight synchronization works unchanged, and it carries no per-node ANNOUNCING/ABORTED sentinels. struct _timeout grows a wheel-only flags field (which wheel tier a timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist and heap builds are unaffected. The backend is EXPERIMENTAL. Two known limitations, inherent to the wheel algorithm (not the abstraction): - No same-tick firing-order guarantee (sifted timeouts are prepended), so the timeout_order test does not apply. - next_timeout() never exceeds 32 ticks because a sift is always pending, so the wheel wakes a tickless-idle CPU at least every 32 ticks. This fails tests that assert zero spurious idle wakeups (tests/kernel/context cpu_idle / timer_interrupts) and is a power regression versus the dlist and heap backends. Verified the timeout-functional suites (timer_api, timeout, timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86, x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected. The timer-wheel data structure, bucketing scheme and sift algorithm are the work of Peter Mitsis (PR #108339), re-homed here behind the timeout backend interface. Co-authored-by: Peter Mitsis <peter.mitsis@intel.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
* shall it call any other subsystem while holding this lock. Code outside this
* file takes it through sys_clock_lock()/sys_clock_unlock().
*/
static struct k_spinlock timeout_lock;
/* Ticks left to process in the currently-executing sys_clock_announce() */
static uint32_t announce_remaining;
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
/* CPU id currently inside sys_clock_announce_locked()'s firing loop, or -1
* when no CPU is. The SMP early-return below ensures at most one CPU is in
* the loop at a time, so a single int suffices. Used by code that needs to
* know whether *this* CPU is at a tick edge (announcing) versus somewhere
* within a tick (any other context, including running on a CPU while a
* different CPU is the announcer).
*/
static int announcing_cpu = -1;
static inline bool this_cpu_announcing(void)
{
return announcing_cpu == CPU_ID;
}
static inline bool any_cpu_announcing(void)
{
return announcing_cpu != -1;
}
/* Timeout whose handler is currently being dispatched, or NULL when no
kernel: timeout: tag inflight_timeout with a superseded bit When z_try_abort_timeout() finds the target timeout already popped from the queue and in flight (its handler dispatching), the abort cannot remove anything from the queue. The handler is about to run (or is running) the timeout's callback; the aborter needs a way to tell it "you were aborted, skip your side effects". main carried that signal in the per-timeout dticks=ABORTED sentinel, which any CPU could write under timeout_lock and the handler checked at entry via z_is_timeout_handler_canceled(). Later commits in this series remove the dticks-cancel mechanism, so the signal needs a new home. Encode it in the low bit of the file-local inflight_timeout pointer (struct _timeout is pointer-aligned, so bit 0 is free): inflight_timeout == NULL no handler in flight inflight_timeout == t handler in flight, not superseded inflight_timeout == t | 1 handler in flight, superseded z_try_abort_timeout() sets the bit whenever the target is the in-flight timeout -- on the announcing CPU (same-CPU IRQ that preempted the dispatch, or a stop after a re-arm) and on another CPU racing the handler. A handler with non-idempotent side effects (currently only k_timer's expiry_fn) checks z_timeout_inflight_superseded() at entry and bails if set. Idempotent handlers (z_thread_timeout, work, poll, ...) tolerate the race and don't need the check. The bit is a best-effort signal, not a barrier: an aborter that sets it after the handler has passed its check has no effect (the handler already committed), exactly as the dticks sentinel behaved on main. The -EAGAIN return for the cross-CPU case is a separate mechanism, used only by callers that must wait for the handler to fully complete (e.g. before freeing the timeout's storage); they spin, while best-effort callers ignore -EAGAIN and rely on the bit. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
* handler is in flight. The announcing CPU sets the pointer before
* calling the handler and clears it afterwards; any aborter may set the
* low "superseded" bit (struct _timeout is pointer-aligned so bit 0 is
* free). Accessors below mask the bit.
*/
static struct _timeout *inflight_timeout;
kernel: timeout: tag inflight_timeout with a superseded bit When z_try_abort_timeout() finds the target timeout already popped from the queue and in flight (its handler dispatching), the abort cannot remove anything from the queue. The handler is about to run (or is running) the timeout's callback; the aborter needs a way to tell it "you were aborted, skip your side effects". main carried that signal in the per-timeout dticks=ABORTED sentinel, which any CPU could write under timeout_lock and the handler checked at entry via z_is_timeout_handler_canceled(). Later commits in this series remove the dticks-cancel mechanism, so the signal needs a new home. Encode it in the low bit of the file-local inflight_timeout pointer (struct _timeout is pointer-aligned, so bit 0 is free): inflight_timeout == NULL no handler in flight inflight_timeout == t handler in flight, not superseded inflight_timeout == t | 1 handler in flight, superseded z_try_abort_timeout() sets the bit whenever the target is the in-flight timeout -- on the announcing CPU (same-CPU IRQ that preempted the dispatch, or a stop after a re-arm) and on another CPU racing the handler. A handler with non-idempotent side effects (currently only k_timer's expiry_fn) checks z_timeout_inflight_superseded() at entry and bails if set. Idempotent handlers (z_thread_timeout, work, poll, ...) tolerate the race and don't need the check. The bit is a best-effort signal, not a barrier: an aborter that sets it after the handler has passed its check has no effect (the handler already committed), exactly as the dticks sentinel behaved on main. The -EAGAIN return for the cross-CPU case is a separate mechanism, used only by callers that must wait for the handler to fully complete (e.g. before freeing the timeout's storage); they spin, while best-effort callers ignore -EAGAIN and rely on the bit. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
#define INFLIGHT_SUPERSEDED_BIT 1UL
static inline struct _timeout *inflight_ptr(void)
{
return (struct _timeout *)((uintptr_t)inflight_timeout & ~INFLIGHT_SUPERSEDED_BIT);
}
static inline void inflight_mark_superseded(void)
{
inflight_timeout = (struct _timeout *)((uintptr_t)inflight_timeout |
INFLIGHT_SUPERSEDED_BIT);
}
static uint32_t elapsed(void)
{
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
/*
* While *this* CPU is executing sys_clock_announce_locked()'s firing
* loop, new relative timeouts scheduled from a callback (or from a
* higher-priority ISR that preempted one) are anchored to the currently
* firing tick (curr_tick), so we report 0.
*
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
* On any other CPU the picture is different: we are not at a tick edge,
* and curr_tick is partway through being advanced by the announcing
* CPU's loop. The invariant we want to preserve is
*
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
* T_real - curr_tick = announce_remaining + sys_clock_elapsed()
*
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
* because the driver bumped its internal announced-cycle baseline to
* (curr_tick_initial + N) * CYC_PER_TICK at ISR entry, while the kernel
* has only advanced curr_tick by the K ticks processed so far -- so
* announce_remaining (= N - K) is the residual that must be added to
* sys_clock_elapsed() to get the real-time delta from curr_tick. This
* keeps sys_clock_tick_get() monotonic across the announce window.
*/
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
if (this_cpu_announcing()) {
return 0U;
}
return sys_clock_elapsed() +
(IS_ENABLED(CONFIG_SMP) ? announce_remaining : 0);
}
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
/*
* Backend queue implementation. The selected backend's queue instance and its
* z_timeout_q_*() operations are private to this file: the backend is a header
* included only here, after the shared state above, so those operations can
* reach curr_tick / announce_remaining / inflight_timeout / timeout_lock
* directly. The per-node helpers (z_init_timeout / z_is_inactive_timeout) are
* tree-wide and live in timeout_q.h.
*/
kernel: timeout: add min-heap backend Add a binary min-heap as a selectable timeout backend alongside the sorted delta list, plugging into the backend abstraction rather than forking timeout.c. Pending timeouts are kept in a min-heap keyed on absolute expiry tick; insertion and arbitrary removal are O(log n), which scales better than the delta list's O(n) insertion when many timeouts are pending. The backend is a new kernel/timeout_minheap.h: the heap instance (a min_heap_ref from the previous commit), its comparator and the z_timeout_q_*() operations, included only by timeout.c (after curr_tick) so the whole backend stays private to that translation unit. struct _timeout gains the abs_ticks + heap_handle representation, selected by Kconfig (the delta list's node + dticks is the #else), with the common fn pointer kept as a shared trailing member; the per-node helpers in timeout_q.h grow a matching min-heap variant. The kernel shell thread dump prints the backend's raw scheduling field (dticks or abs_ticks) under #ifdef. Because the shared z_add_timeout() already applies the post-#107452 conditional tick round-up, the heap inherits it: the backend's insert simply stores abs_ticks = curr_tick + dticks. Likewise in-flight handler synchronization (PR #109977) lives in timeout.c, so the heap node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just heap_handle.idx == 0. The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks need 64-bit precision), and uses a fixed-capacity heap (CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error. The min-heap algorithm, struct fields, Kconfig and capacity model are derived from Sayooj K Karun's min-heap timeout subsystem (#106013), reworked here to fit the pluggable backend and the current in-tree in-flight-handler synchronization. Co-authored-by: Sayooj K Karun <sayooj@aerlync.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
#if defined(CONFIG_TIMEOUT_BACKEND_MINHEAP)
#include "timeout_minheap.h"
kernel: timeout: add timer-wheel backend Add a hierarchical timer wheel as a third selectable timeout backend. Timeouts are bucketed by expiry distance: one list per tick for the next 32 ticks ("soon"), one list per 32-tick band for the next ~1024 ("later"), and a sorted overflow list ("distant"). Insertion and removal are O(1) for the common near-future case; every 32 ticks the announce path sifts the current "later" band into "soon" and refills it from "distant". This scales well when many short-lived timeouts are pending. Unlike the dlist and min-heap backends, the wheel does not fit the generic next_gap/advance/pop_due announce primitives: its per-tick advance is a bitmap scan that jumps over empty ticks, and its sift is a time-driven event tied to no single timeout. Rather than contort that (already subtle) state machine, the wheel uses the backend-owned announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and implements z_timeout_q_announce(), which sys_clock_announce_locked() calls in place of the generic loop. Like the other backends the wheel is a single implementation header (kernel/timeout_wheel.h) included only by timeout.c, so its state, its operations and that announce loop reach the shared state (curr_tick, announce_remaining, inflight_timeout) directly; nothing extra needs exposing. The SMP re-entry guard, announcing_cpu and the reprogram remain in timeout.c. The wheel fires handlers through the same inflight_timeout dance, so the post-#109977 abort/in-flight synchronization works unchanged, and it carries no per-node ANNOUNCING/ABORTED sentinels. struct _timeout grows a wheel-only flags field (which wheel tier a timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist and heap builds are unaffected. The backend is EXPERIMENTAL. Two known limitations, inherent to the wheel algorithm (not the abstraction): - No same-tick firing-order guarantee (sifted timeouts are prepended), so the timeout_order test does not apply. - next_timeout() never exceeds 32 ticks because a sift is always pending, so the wheel wakes a tickless-idle CPU at least every 32 ticks. This fails tests that assert zero spurious idle wakeups (tests/kernel/context cpu_idle / timer_interrupts) and is a power regression versus the dlist and heap backends. Verified the timeout-functional suites (timer_api, timeout, timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86, x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected. The timer-wheel data structure, bucketing scheme and sift algorithm are the work of Peter Mitsis (PR #108339), re-homed here behind the timeout backend interface. Co-authored-by: Peter Mitsis <peter.mitsis@intel.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
#elif defined(CONFIG_TIMEOUT_BACKEND_WHEEL)
#include "timeout_wheel.h"
kernel: timeout: add bucketed delta-list backend Add a single-level bucketed delta list as a fourth selectable timeout backend: a simpler relative of the timer wheel that keeps the wheel's O(1) near-future insertion without its second tier, sift/defer machinery, or tickless-idle penalty. Timeouts expiring within the next CONFIG_TIMEOUT_BUCKET_LISTS ticks go into a per-tick bucket list (O(1) insert and remove), tracked by an occupancy bitmap so "ticks until next expiry" is O(1). Everything beyond goes into one overflow list sorted by absolute expiry (O(n) insert, O(1) remove, no delta fix-up). As curr_tick crosses the bucket window, overflow entries that have come within range migrate into buckets. Like the other backends it is a single implementation header (kernel/timeout_bucket.h) included only by timeout.c. Two properties distinguish it from the other non-default backends, both confirmed by running the full kernel timer/context/common suites with the backend forced on (qemu_x86, x86_64, cortex_a53 SMP, riscv64): - It fits the generic next_gap/advance/pop_due announce loop -- no backend-owned announce. Migration is on demand against the actual next event, so an idle CPU with a distant timeout sleeps straight to it: it passes tests/kernel/context cpu_idle / timer_interrupts, which the wheel fails (the wheel wakes at least every 32 ticks). - Same-tick firing order stays FIFO (bucket and overflow inserts append; migration preserves order), so tests/kernel/common's timeout_order passes, unlike the min-heap. The per-node representation is the delta list's node + dticks with no extra field: a bucket entry stores its bucket index (< BUCKET_LISTS) in dticks, an overflow entry its absolute expiry (>= BUCKET_LISTS); the ranges never overlap, so it shares the dlist per-node helpers. Absolute expiry needs 64-bit ticks, so the backend depends on TIMEOUT_64BIT. It inherits the shared z_add_timeout round-up and the inflight_timeout synchronization like the other backends. EXPERIMENTAL. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 22:00:48 -04:00
#elif defined(CONFIG_TIMEOUT_BACKEND_BUCKET)
#include "timeout_bucket.h"
#elif defined(CONFIG_TIMEOUT_BACKEND_SKIPLIST)
#include "timeout_skiplist.h"
kernel: timeout: add min-heap backend Add a binary min-heap as a selectable timeout backend alongside the sorted delta list, plugging into the backend abstraction rather than forking timeout.c. Pending timeouts are kept in a min-heap keyed on absolute expiry tick; insertion and arbitrary removal are O(log n), which scales better than the delta list's O(n) insertion when many timeouts are pending. The backend is a new kernel/timeout_minheap.h: the heap instance (a min_heap_ref from the previous commit), its comparator and the z_timeout_q_*() operations, included only by timeout.c (after curr_tick) so the whole backend stays private to that translation unit. struct _timeout gains the abs_ticks + heap_handle representation, selected by Kconfig (the delta list's node + dticks is the #else), with the common fn pointer kept as a shared trailing member; the per-node helpers in timeout_q.h grow a matching min-heap variant. The kernel shell thread dump prints the backend's raw scheduling field (dticks or abs_ticks) under #ifdef. Because the shared z_add_timeout() already applies the post-#107452 conditional tick round-up, the heap inherits it: the backend's insert simply stores abs_ticks = curr_tick + dticks. Likewise in-flight handler synchronization (PR #109977) lives in timeout.c, so the heap node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just heap_handle.idx == 0. The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks need 64-bit precision), and uses a fixed-capacity heap (CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error. The min-heap algorithm, struct fields, Kconfig and capacity model are derived from Sayooj K Karun's min-heap timeout subsystem (#106013), reworked here to fit the pluggable backend and the current in-tree in-flight-handler synchronization. Co-authored-by: Sayooj K Karun <sayooj@aerlync.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
#else /* CONFIG_TIMEOUT_BACKEND_DLIST */
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
#include "timeout_list.h"
kernel: timeout: add min-heap backend Add a binary min-heap as a selectable timeout backend alongside the sorted delta list, plugging into the backend abstraction rather than forking timeout.c. Pending timeouts are kept in a min-heap keyed on absolute expiry tick; insertion and arbitrary removal are O(log n), which scales better than the delta list's O(n) insertion when many timeouts are pending. The backend is a new kernel/timeout_minheap.h: the heap instance (a min_heap_ref from the previous commit), its comparator and the z_timeout_q_*() operations, included only by timeout.c (after curr_tick) so the whole backend stays private to that translation unit. struct _timeout gains the abs_ticks + heap_handle representation, selected by Kconfig (the delta list's node + dticks is the #else), with the common fn pointer kept as a shared trailing member; the per-node helpers in timeout_q.h grow a matching min-heap variant. The kernel shell thread dump prints the backend's raw scheduling field (dticks or abs_ticks) under #ifdef. Because the shared z_add_timeout() already applies the post-#107452 conditional tick round-up, the heap inherits it: the backend's insert simply stores abs_ticks = curr_tick + dticks. Likewise in-flight handler synchronization (PR #109977) lives in timeout.c, so the heap node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just heap_handle.idx == 0. The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks need 64-bit precision), and uses a fixed-capacity heap (CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error. The min-heap algorithm, struct fields, Kconfig and capacity model are derived from Sayooj K Karun's min-heap timeout subsystem (#106013), reworked here to fit the pluggable backend and the current in-tree in-flight-handler synchronization. Co-authored-by: Sayooj K Karun <sayooj@aerlync.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 18:20:59 -04:00
#endif
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
/*
* Ticks the driver may wait before the next sys_clock_announce(). The
* announce-range cap lives here, in the backend-independent front end, so no
* backend has to reproduce it: the backend only reports the delta to its
* earliest pending timeout via z_timeout_q_next_expiry().
*/
static uint32_t next_timeout(uint32_t ticks_elapsed)
{
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
k_ticks_t next = z_timeout_q_next_expiry();
uint32_t dticks;
/*
* sys_clock_announce() reports the ticks elapsed since the previous
* announce, so the gap between two announces must not exceed
* SYS_CLOCK_MAX_WAIT for the announced count to fit. The budget left
* from now on is SYS_CLOCK_MAX_WAIT - ticks_elapsed; if it is already
* spent ask for an announce right away.
*/
if (ticks_elapsed >= SYS_CLOCK_MAX_WAIT) {
return 0;
}
/*
* No deadline, or one further out than can be scheduled in a single
* step: wait the capped budget and re-evaluate at the next announce.
* Testing next (which may be 64-bit) against the cap keeps the
* remaining arithmetic in 32 bits. The empty case still returns the
* budget so a driver keeps waking to maintain uptime.
*/
if (next == K_TICKS_FOREVER || next >= SYS_CLOCK_MAX_WAIT) {
return SYS_CLOCK_MAX_WAIT - ticks_elapsed;
}
/* Otherwise wait until the timeout, relative to now (0 if due). */
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
dticks = (uint32_t)next;
return (dticks > ticks_elapsed) ? (dticks - ticks_elapsed) : 0;
}
/*
* Reprogram the timer for the next pending timeout, or, when nothing is
* pending and CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE tolerates a drifting uptime, tell
* the driver its clock is unused so it may stop. Only meaningful where the
* timeout queue may have just drained (abort, end of announce); the add path
* always has a pending timeout and calls sys_clock_set_timeout() directly.
*/
static void reprogram_next(uint32_t ticks_elapsed)
{
if (IS_ENABLED(CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE) &&
z_timeout_q_next_expiry() == K_TICKS_FOREVER) {
sys_clock_no_timeout();
} else {
sys_clock_set_timeout(next_timeout(ticks_elapsed), false);
}
}
k_ticks_t z_add_timeout(struct _timeout *to, _timeout_func_t fn, k_timeout_t timeout)
{
k_ticks_t ticks = 0;
if (K_TIMEOUT_EQ(timeout, K_FOREVER)) {
return 0;
}
#ifdef CONFIG_KERNEL_COHERENCE
__ASSERT_NO_MSG(sys_cache_is_mem_coherent(to));
#endif /* CONFIG_KERNEL_COHERENCE */
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
__ASSERT(z_is_inactive_timeout(to), "");
to->fn = fn;
K_SPINLOCK(&timeout_lock) {
uint32_t ticks_elapsed = 0;
bool has_elapsed = false;
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
k_ticks_t dticks;
if (Z_IS_TIMEOUT_RELATIVE(timeout)) {
ticks_elapsed = elapsed();
has_elapsed = true;
kernel: timeout: make z_add_timeout round-up conditional on announce z_add_timeout() has always added one tick to the incoming tick count, as a conservative round-up so that a request issued partway through a tick still waits for "at least N full ticks" before the fire. This round-up is correct in the general case, but *wrong* when the call happens from within sys_clock_announce_locked() -- i.e. from a timer expiration callback running at the tick-processing boundary where elapsed() already returns 0. In that context there is no fractional tick to compensate for, and the round-up simply makes every scheduled timeout one tick late. Two in-tree callers were already compensating for this caller-side: * The k_timer periodic reschedule path in z_timer_expiration_handler always runs from inside sys_clock_announce_locked() and subtracted 1 from the period before calling z_add_timeout(). Under CONFIG_TIMEOUT_64BIT, the same path additionally added +1 inside K_TIMEOUT_ABS_TICKS() to undo a related round-down. * z_time_slice_reset() armed the slice timer with K_TICKS(slice_size - 1) so the resulting fire would land at exactly slice_size ticks. This one is reachable from both thread context (the +1 cancels the -1) and from announce context via update_cache() during a ready-thread wakeup (the +1 isn't applied, leaving the slice short by one tick). The latter is what actually trips tests/kernel/tickless/tickless_concept on every platform once the conditional below lands without dropping these workarounds. All three are symptoms of the same root cause. Handle it at the source: make the +1 conditional on announce_remaining == 0. When scheduling from the timer ISR, we are already at a tick boundary by construction, so no round-up is needed and periodic timers now reschedule at exact period intervals without any caller-side compensation. Drop the -1 in z_timer_expiration_handler's period path, the +1 in its 64-bit absolute-reschedule companion, and the -1 in z_time_slice_reset(), since all three existed solely to paper over this mismatch. This change only affects timeouts scheduled from announce context (periodic k_timers rescheduling themselves, callbacks starting new timers, slice-timer rearm during ready-thread wakeup). All other callers -- k_sleep(), z_abort_timeout(), initial k_timer_start() from a thread, k_sched_time_slice_set() -- continue to use the +1 round-up. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-17 00:36:30 -04:00
/*
* In the general case, "now" may be anywhere within
* the current tick. Rounding up by one tick guarantees
* "at least N ticks" semantics -- otherwise a request
* made partway through a tick would fire on the next
* tick edge, yielding less than N full ticks.
*
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
* The one moment we know *this* CPU is at a tick edge
* is while it is processing timeouts inside its own
* sys_clock_announce_locked() loop. Periodic timers
* rely on this when rescheduling themselves from the
* timer ISR: the round-up would otherwise accumulate
* and make every period one tick late. The check is
* per-CPU: a thread on another CPU running while we
* announce is *not* at a tick edge and still needs
* the round-up.
kernel: timeout: make z_add_timeout round-up conditional on announce z_add_timeout() has always added one tick to the incoming tick count, as a conservative round-up so that a request issued partway through a tick still waits for "at least N full ticks" before the fire. This round-up is correct in the general case, but *wrong* when the call happens from within sys_clock_announce_locked() -- i.e. from a timer expiration callback running at the tick-processing boundary where elapsed() already returns 0. In that context there is no fractional tick to compensate for, and the round-up simply makes every scheduled timeout one tick late. Two in-tree callers were already compensating for this caller-side: * The k_timer periodic reschedule path in z_timer_expiration_handler always runs from inside sys_clock_announce_locked() and subtracted 1 from the period before calling z_add_timeout(). Under CONFIG_TIMEOUT_64BIT, the same path additionally added +1 inside K_TIMEOUT_ABS_TICKS() to undo a related round-down. * z_time_slice_reset() armed the slice timer with K_TICKS(slice_size - 1) so the resulting fire would land at exactly slice_size ticks. This one is reachable from both thread context (the +1 cancels the -1) and from announce context via update_cache() during a ready-thread wakeup (the +1 isn't applied, leaving the slice short by one tick). The latter is what actually trips tests/kernel/tickless/tickless_concept on every platform once the conditional below lands without dropping these workarounds. All three are symptoms of the same root cause. Handle it at the source: make the +1 conditional on announce_remaining == 0. When scheduling from the timer ISR, we are already at a tick boundary by construction, so no round-up is needed and periodic timers now reschedule at exact period intervals without any caller-side compensation. Drop the -1 in z_timer_expiration_handler's period path, the +1 in its 64-bit absolute-reschedule companion, and the -1 in z_time_slice_reset(), since all three existed solely to paper over this mismatch. This change only affects timeouts scheduled from announce context (periodic k_timers rescheduling themselves, callbacks starting new timers, slice-timer rearm during ready-thread wakeup). All other callers -- k_sleep(), z_abort_timeout(), initial k_timer_start() from a thread, k_sched_time_slice_set() -- continue to use the +1 round-up. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-17 00:36:30 -04:00
*/
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
dticks = timeout.ticks + ticks_elapsed +
(this_cpu_announcing() ? 0 : 1);
ticks = curr_tick + dticks;
} else {
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
dticks = Z_TICK_ABS(timeout.ticks) - curr_tick;
dticks = max(1, dticks);
ticks = timeout.ticks;
}
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
if (z_timeout_q_insert(to, dticks) && !any_cpu_announcing()) {
if (!has_elapsed) {
/* In case of absolute timeout that is first to expire
* elapsed need to be read from the system clock.
*/
ticks_elapsed = elapsed();
}
sys_clock_set_timeout(next_timeout(ticks_elapsed), false);
}
}
return ticks;
}
int z_try_abort_timeout(struct _timeout *to)
{
int ret = -EINVAL;
K_SPINLOCK(&timeout_lock) {
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
if (!z_is_inactive_timeout(to)) {
bool was_first = z_timeout_q_remove(to);
ret = 0;
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
if (was_first) {
reprogram_next(elapsed());
}
kernel: timeout: tag inflight_timeout with a superseded bit When z_try_abort_timeout() finds the target timeout already popped from the queue and in flight (its handler dispatching), the abort cannot remove anything from the queue. The handler is about to run (or is running) the timeout's callback; the aborter needs a way to tell it "you were aborted, skip your side effects". main carried that signal in the per-timeout dticks=ABORTED sentinel, which any CPU could write under timeout_lock and the handler checked at entry via z_is_timeout_handler_canceled(). Later commits in this series remove the dticks-cancel mechanism, so the signal needs a new home. Encode it in the low bit of the file-local inflight_timeout pointer (struct _timeout is pointer-aligned, so bit 0 is free): inflight_timeout == NULL no handler in flight inflight_timeout == t handler in flight, not superseded inflight_timeout == t | 1 handler in flight, superseded z_try_abort_timeout() sets the bit whenever the target is the in-flight timeout -- on the announcing CPU (same-CPU IRQ that preempted the dispatch, or a stop after a re-arm) and on another CPU racing the handler. A handler with non-idempotent side effects (currently only k_timer's expiry_fn) checks z_timeout_inflight_superseded() at entry and bails if set. Idempotent handlers (z_thread_timeout, work, poll, ...) tolerate the race and don't need the check. The bit is a best-effort signal, not a barrier: an aborter that sets it after the handler has passed its check has no effect (the handler already committed), exactly as the dticks sentinel behaved on main. The -EAGAIN return for the cross-CPU case is a separate mechanism, used only by callers that must wait for the handler to fully complete (e.g. before freeing the timeout's storage); they spin, while best-effort callers ignore -EAGAIN and rely on the bit. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
} else if (IS_ENABLED(CONFIG_SMP) && inflight_ptr() == to &&
!this_cpu_announcing()) {
/* Handler in flight on another CPU. Free-safety
* callers retry on -EAGAIN to wait it out; others
* rely on the superseded mark below and don't.
*/
ret = -EAGAIN;
}
/* Record that the in-flight timeout has been aborted, so a
* handler that checks (z_timeout_inflight_superseded) can
* tell its dispatch was overtaken.
*/
if (inflight_ptr() == to) {
inflight_mark_superseded();
}
}
if (IS_ENABLED(CONFIG_SMP) && ret == -EAGAIN) {
arch_spin_relax();
}
return ret;
}
kernel: timeout: tag inflight_timeout with a superseded bit When z_try_abort_timeout() finds the target timeout already popped from the queue and in flight (its handler dispatching), the abort cannot remove anything from the queue. The handler is about to run (or is running) the timeout's callback; the aborter needs a way to tell it "you were aborted, skip your side effects". main carried that signal in the per-timeout dticks=ABORTED sentinel, which any CPU could write under timeout_lock and the handler checked at entry via z_is_timeout_handler_canceled(). Later commits in this series remove the dticks-cancel mechanism, so the signal needs a new home. Encode it in the low bit of the file-local inflight_timeout pointer (struct _timeout is pointer-aligned, so bit 0 is free): inflight_timeout == NULL no handler in flight inflight_timeout == t handler in flight, not superseded inflight_timeout == t | 1 handler in flight, superseded z_try_abort_timeout() sets the bit whenever the target is the in-flight timeout -- on the announcing CPU (same-CPU IRQ that preempted the dispatch, or a stop after a re-arm) and on another CPU racing the handler. A handler with non-idempotent side effects (currently only k_timer's expiry_fn) checks z_timeout_inflight_superseded() at entry and bails if set. Idempotent handlers (z_thread_timeout, work, poll, ...) tolerate the race and don't need the check. The bit is a best-effort signal, not a barrier: an aborter that sets it after the handler has passed its check has no effect (the handler already committed), exactly as the dticks sentinel behaved on main. The -EAGAIN return for the cross-CPU case is a separate mechanism, used only by callers that must wait for the handler to fully complete (e.g. before freeing the timeout's storage); they spin, while best-effort callers ignore -EAGAIN and rely on the bit. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-28 19:39:37 -04:00
bool z_timeout_inflight_superseded(const struct _timeout *to)
{
bool superseded = false;
K_SPINLOCK(&timeout_lock) {
superseded = inflight_timeout ==
(struct _timeout *)((uintptr_t)to | INFLIGHT_SUPERSEDED_BIT);
}
return superseded;
}
k_ticks_t z_timeout_remaining(const struct _timeout *timeout)
{
k_ticks_t ticks = 0;
K_SPINLOCK(&timeout_lock) {
if (!z_is_inactive_timeout(timeout)) {
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
ticks = z_timeout_q_remainder(timeout) - elapsed();
}
}
return ticks;
}
EXPORT_SYMBOL(z_timeout_remaining);
k_ticks_t z_timeout_expires(const struct _timeout *timeout)
{
k_ticks_t ticks = 0;
K_SPINLOCK(&timeout_lock) {
ticks = curr_tick;
if (!z_is_inactive_timeout(timeout)) {
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
ticks += z_timeout_q_remainder(timeout);
}
}
return ticks;
}
EXPORT_SYMBOL(z_timeout_expires);
uint32_t z_get_next_timeout_expiry(void)
{
uint32_t ret = (uint32_t)K_TICKS_FOREVER;
K_SPINLOCK(&timeout_lock) {
/*
* Same decision as reprogram_next(). Sloppy idle lets the
* uptime drift, so an empty list means nothing to wake up for
* and that is reported as K_TICKS_FOREVER. Otherwise the
* answer is a wait: either a real deadline, or the synthetic
* one that keeps the announce range covered.
*/
if (IS_ENABLED(CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE) &&
z_timeout_q_next_expiry() == K_TICKS_FOREVER) {
ret = (uint32_t)K_TICKS_FOREVER;
} else {
ret = next_timeout(elapsed());
}
}
return ret;
}
void sys_clock_announce_locked(uint32_t ticks, k_spinlock_key_t key)
{
/* We release the lock around the callbacks below, so on SMP
* systems someone might be already running the loop. Don't
* race (which will cause parallel execution of "sequential"
* timeouts and confuse apps), just increment the tick count
* and return.
*/
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
if (IS_ENABLED(CONFIG_SMP) && any_cpu_announcing()) {
announce_remaining += ticks;
k_spin_unlock(&timeout_lock, key);
return;
}
announce_remaining = ticks;
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
announcing_cpu = CPU_ID;
kernel: timeout: add timer-wheel backend Add a hierarchical timer wheel as a third selectable timeout backend. Timeouts are bucketed by expiry distance: one list per tick for the next 32 ticks ("soon"), one list per 32-tick band for the next ~1024 ("later"), and a sorted overflow list ("distant"). Insertion and removal are O(1) for the common near-future case; every 32 ticks the announce path sifts the current "later" band into "soon" and refills it from "distant". This scales well when many short-lived timeouts are pending. Unlike the dlist and min-heap backends, the wheel does not fit the generic next_gap/advance/pop_due announce primitives: its per-tick advance is a bitmap scan that jumps over empty ticks, and its sift is a time-driven event tied to no single timeout. Rather than contort that (already subtle) state machine, the wheel uses the backend-owned announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and implements z_timeout_q_announce(), which sys_clock_announce_locked() calls in place of the generic loop. Like the other backends the wheel is a single implementation header (kernel/timeout_wheel.h) included only by timeout.c, so its state, its operations and that announce loop reach the shared state (curr_tick, announce_remaining, inflight_timeout) directly; nothing extra needs exposing. The SMP re-entry guard, announcing_cpu and the reprogram remain in timeout.c. The wheel fires handlers through the same inflight_timeout dance, so the post-#109977 abort/in-flight synchronization works unchanged, and it carries no per-node ANNOUNCING/ABORTED sentinels. struct _timeout grows a wheel-only flags field (which wheel tier a timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist and heap builds are unaffected. The backend is EXPERIMENTAL. Two known limitations, inherent to the wheel algorithm (not the abstraction): - No same-tick firing-order guarantee (sifted timeouts are prepended), so the timeout_order test does not apply. - next_timeout() never exceeds 32 ticks because a sift is always pending, so the wheel wakes a tickless-idle CPU at least every 32 ticks. This fails tests that assert zero spurious idle wakeups (tests/kernel/context cpu_idle / timer_interrupts) and is a power regression versus the dlist and heap backends. Verified the timeout-functional suites (timer_api, timeout, timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86, x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected. The timer-wheel data structure, bucketing scheme and sift algorithm are the work of Peter Mitsis (PR #108339), re-homed here behind the timeout backend interface. Co-authored-by: Peter Mitsis <peter.mitsis@intel.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
#ifdef _TIMEOUT_BACKEND_OWNS_ANNOUNCE
/* The backend (timer wheel) runs the firing loop itself: its per-tick
* advance is a bitmap jump over empty ticks and its sift is a
* time-driven event with no single timeout to drive a generic loop.
* It advances curr_tick / announce_remaining and fires handlers via
* the same inflight_timeout dance as below.
*/
key = z_timeout_q_announce(key);
#else
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
while (announce_remaining > 0) {
struct _timeout *t;
int32_t dt = z_timeout_q_next_gap();
if ((uint32_t)dt > announce_remaining) {
/* Next event is past this announce window: advance the
* backend by the residual ticks without firing anything.
*/
dt = (int32_t)announce_remaining;
}
/* Advance curr_tick and decrement announce_remaining together
* under the lock so non-announcing CPUs observe a consistent
* (curr_tick + announce_remaining + sys_clock_elapsed()) ==
* T_real even while we drop the lock around handlers. The "we
* are announcing" state is carried by announcing_cpu, so
* announce_remaining reaching 0 mid-loop is harmless: same-tick
* handlers' z_add_timeout() still anchors via
* this_cpu_announcing(), and another CPU's announce still folds
* into ours via any_cpu_announcing() in the SMP early-return.
kernel: timeout: keep announce_remaining stable across same-tick group When sys_clock_announce_locked() processes a tick that has multiple timeouts queued for it, the timeout queue stores the second and subsequent ones with dticks == 0 (relative to the first). The original loop fired each in turn and decremented announce_remaining at the bottom of every iteration: announce_remaining -= dt; For the first timeout in a same-tick group, dt is the cumulative tick delta that brought us to that tick. After that subtraction announce_remaining can drop to zero, even though there are still same-tick callbacks queued for the loop to fire. Each subsequent same-tick callback then runs while announce_remaining == 0, which breaks two invariants the rest of the kernel relies on: * The SMP early-return at the top of sys_clock_announce_locked(): if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) { announce_remaining += ticks; k_spin_unlock(&timeout_lock, key); return; } is meant to detect that another CPU is already inside the loop and fold the new ticks into the ongoing announce. The lock is released around each callback, so during a same-tick callback another CPU can grab the lock, see announce_remaining == 0, miss the early return, set announce_remaining = ticks of its own, and start walking the queue in parallel with the original announcer -- the exact race the early return is supposed to prevent. * The elapsed() helper: return announce_remaining == 0 ? sys_clock_elapsed() : 0U; is meant to return 0 for any z_add_timeout() / z_abort_timeout() call that happens from inside a tick-processing callback, so that timeouts scheduled from such a callback are anchored to the currently-firing tick rather than to a fresh sys_clock_elapsed() reading. With announce_remaining == 0 mid-loop, two callbacks on the same tick observe inconsistent semantics: the first one (that saw announce_remaining > 0) gets dticks anchored to the firing tick, while subsequent ones (seeing 0) get dticks computed against a fresh elapsed() reading and end up off by one tick. Two periodic timers that happen to fire on the same tick will therefore permanently drift apart by one tick going forward. Restructure the loop so the announce_remaining decrement happens once per tick rather than once per timeout: the outer while drives forward across distinct ticks, and an inner do-while drains all timeouts queued on the current tick before announce_remaining is updated. announce_remaining now stays at its pre-tick value for the entire same-tick group, which both the SMP early return and elapsed() correctly observe as non-zero. remove_timeout()'s dticks propagation is also unnecessary in this loop because curr_tick is advanced by t->dticks before the timeout is unlinked, which keeps the next item's stored dticks valid relative to the new curr_tick. sys_dlist_remove() on its own is sufficient. Loop structure suggested by Peter Mitsis. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
*/
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
z_timeout_q_advance(dt);
curr_tick += dt;
announce_remaining -= dt;
kernel: timeout: keep announce_remaining stable across same-tick group When sys_clock_announce_locked() processes a tick that has multiple timeouts queued for it, the timeout queue stores the second and subsequent ones with dticks == 0 (relative to the first). The original loop fired each in turn and decremented announce_remaining at the bottom of every iteration: announce_remaining -= dt; For the first timeout in a same-tick group, dt is the cumulative tick delta that brought us to that tick. After that subtraction announce_remaining can drop to zero, even though there are still same-tick callbacks queued for the loop to fire. Each subsequent same-tick callback then runs while announce_remaining == 0, which breaks two invariants the rest of the kernel relies on: * The SMP early-return at the top of sys_clock_announce_locked(): if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) { announce_remaining += ticks; k_spin_unlock(&timeout_lock, key); return; } is meant to detect that another CPU is already inside the loop and fold the new ticks into the ongoing announce. The lock is released around each callback, so during a same-tick callback another CPU can grab the lock, see announce_remaining == 0, miss the early return, set announce_remaining = ticks of its own, and start walking the queue in parallel with the original announcer -- the exact race the early return is supposed to prevent. * The elapsed() helper: return announce_remaining == 0 ? sys_clock_elapsed() : 0U; is meant to return 0 for any z_add_timeout() / z_abort_timeout() call that happens from inside a tick-processing callback, so that timeouts scheduled from such a callback are anchored to the currently-firing tick rather than to a fresh sys_clock_elapsed() reading. With announce_remaining == 0 mid-loop, two callbacks on the same tick observe inconsistent semantics: the first one (that saw announce_remaining > 0) gets dticks anchored to the firing tick, while subsequent ones (seeing 0) get dticks computed against a fresh elapsed() reading and end up off by one tick. Two periodic timers that happen to fire on the same tick will therefore permanently drift apart by one tick going forward. Restructure the loop so the announce_remaining decrement happens once per tick rather than once per timeout: the outer while drives forward across distinct ticks, and an inner do-while drains all timeouts queued on the current tick before announce_remaining is updated. announce_remaining now stays at its pre-tick value for the entire same-tick group, which both the SMP early return and elapsed() correctly observe as non-zero. remove_timeout()'s dticks propagation is also unnecessary in this loop because curr_tick is advanced by t->dticks before the timeout is unlinked, which keeps the next item's stored dticks valid relative to the new curr_tick. sys_dlist_remove() on its own is sufficient. Loop structure suggested by Peter Mitsis. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
/* Drain everything due at the new curr_tick. */
while ((t = z_timeout_q_pop_due()) != NULL) {
_timeout_func_t handler = t->fn;
kernel: timeout: keep announce_remaining stable across same-tick group When sys_clock_announce_locked() processes a tick that has multiple timeouts queued for it, the timeout queue stores the second and subsequent ones with dticks == 0 (relative to the first). The original loop fired each in turn and decremented announce_remaining at the bottom of every iteration: announce_remaining -= dt; For the first timeout in a same-tick group, dt is the cumulative tick delta that brought us to that tick. After that subtraction announce_remaining can drop to zero, even though there are still same-tick callbacks queued for the loop to fire. Each subsequent same-tick callback then runs while announce_remaining == 0, which breaks two invariants the rest of the kernel relies on: * The SMP early-return at the top of sys_clock_announce_locked(): if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) { announce_remaining += ticks; k_spin_unlock(&timeout_lock, key); return; } is meant to detect that another CPU is already inside the loop and fold the new ticks into the ongoing announce. The lock is released around each callback, so during a same-tick callback another CPU can grab the lock, see announce_remaining == 0, miss the early return, set announce_remaining = ticks of its own, and start walking the queue in parallel with the original announcer -- the exact race the early return is supposed to prevent. * The elapsed() helper: return announce_remaining == 0 ? sys_clock_elapsed() : 0U; is meant to return 0 for any z_add_timeout() / z_abort_timeout() call that happens from inside a tick-processing callback, so that timeouts scheduled from such a callback are anchored to the currently-firing tick rather than to a fresh sys_clock_elapsed() reading. With announce_remaining == 0 mid-loop, two callbacks on the same tick observe inconsistent semantics: the first one (that saw announce_remaining > 0) gets dticks anchored to the firing tick, while subsequent ones (seeing 0) get dticks computed against a fresh elapsed() reading and end up off by one tick. Two periodic timers that happen to fire on the same tick will therefore permanently drift apart by one tick going forward. Restructure the loop so the announce_remaining decrement happens once per tick rather than once per timeout: the outer while drives forward across distinct ticks, and an inner do-while drains all timeouts queued on the current tick before announce_remaining is updated. announce_remaining now stays at its pre-tick value for the entire same-tick group, which both the SMP early return and elapsed() correctly observe as non-zero. remove_timeout()'s dticks propagation is also unnecessary in this loop because curr_tick is advanced by t->dticks before the timeout is unlinked, which keeps the next item's stored dticks valid relative to the new curr_tick. sys_dlist_remove() on its own is sufficient. Loop structure suggested by Peter Mitsis. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-04-29 16:24:46 -04:00
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
inflight_timeout = t;
kernel: timeout: factor delta-list queue into a pluggable backend The timeout queue data structure and the front-end logic in timeout.c are entangled through first()/next()/remove_timeout()/next_timeout() and the open-coded insertion and announce loops. This makes it hard to offer alternative queue implementations (min-heap, timer wheel) without either forking timeout.c or threading #ifdefs through its trickiest code, the announce path in particular. Introduce a backend interface and move the sorted delta list behind it. The split mirrors how the queue is actually used: - The per-node helpers z_init_timeout() / z_is_inactive_timeout() operate on a single struct _timeout and are needed tree-wide (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and depend only on struct _timeout's fields. - The queue itself (its instance and the z_timeout_q_*() operations) is used only by timeout.c, so it becomes a backend implementation header (kernel/timeout_list.h) that timeout.c includes directly, after its shared state. The header is private to that translation unit; no other file sees the queue or its instance, so the instance is plain static and needs no extern. timeout.c keeps all shared state (curr_tick, announce_remaining, announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API translation, the in-flight handler synchronization, and a single announce loop shared by every future backend. The backend exposes: - z_timeout_q_insert / _remove add and arbitrary remove - z_timeout_q_remainder / _next_timeout queries - z_timeout_q_next_gap / _advance / _pop_due announce-loop primitives The announce loop is recast in terms of the last three primitives: find the gap to the next event, advance the backend (and curr_tick) to it, then drain whatever is due. For the delta list this is behaviorally identical to the previous open-coded loop: the residual advance reuses the same head->dticks fixup, same-tick ties drain via successive pop_due() calls, and announce_remaining is still decremented per tick group while announcing_cpu carries the announcing state. So the front end no longer reaches into the queue representation, three checks are made backend-neutral: z_add_timeout()'s assert and z_timer_expiration_handler()'s "restarted?" test now use z_is_inactive_timeout() instead of sys_dnode_is_linked(), and Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than naming .node/.dticks. struct _timeout is otherwise untouched; no sentinel or per-node state changes. The Kconfig backend choice and the divergent node layouts are deferred until a second backend lands; with only the delta list present they would be churn with no benefit. Verified no behavioral change: tests/kernel/timer/{timer_api,timeout}, tests/kernel/common, and tests/kernel/sleep pass on qemu_x86, qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the SMP announcing_cpu path and 64-bit dticks). Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-19 17:49:31 -04:00
k_spin_unlock(&timeout_lock, key);
handler(t);
key = k_spin_lock(&timeout_lock);
inflight_timeout = NULL;
}
}
kernel: timeout: add timer-wheel backend Add a hierarchical timer wheel as a third selectable timeout backend. Timeouts are bucketed by expiry distance: one list per tick for the next 32 ticks ("soon"), one list per 32-tick band for the next ~1024 ("later"), and a sorted overflow list ("distant"). Insertion and removal are O(1) for the common near-future case; every 32 ticks the announce path sifts the current "later" band into "soon" and refills it from "distant". This scales well when many short-lived timeouts are pending. Unlike the dlist and min-heap backends, the wheel does not fit the generic next_gap/advance/pop_due announce primitives: its per-tick advance is a bitmap scan that jumps over empty ticks, and its sift is a time-driven event tied to no single timeout. Rather than contort that (already subtle) state machine, the wheel uses the backend-owned announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and implements z_timeout_q_announce(), which sys_clock_announce_locked() calls in place of the generic loop. Like the other backends the wheel is a single implementation header (kernel/timeout_wheel.h) included only by timeout.c, so its state, its operations and that announce loop reach the shared state (curr_tick, announce_remaining, inflight_timeout) directly; nothing extra needs exposing. The SMP re-entry guard, announcing_cpu and the reprogram remain in timeout.c. The wheel fires handlers through the same inflight_timeout dance, so the post-#109977 abort/in-flight synchronization works unchanged, and it carries no per-node ANNOUNCING/ABORTED sentinels. struct _timeout grows a wheel-only flags field (which wheel tier a timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist and heap builds are unaffected. The backend is EXPERIMENTAL. Two known limitations, inherent to the wheel algorithm (not the abstraction): - No same-tick firing-order guarantee (sifted timeouts are prepended), so the timeout_order test does not apply. - next_timeout() never exceeds 32 ticks because a sift is always pending, so the wheel wakes a tickless-idle CPU at least every 32 ticks. This fails tests that assert zero spurious idle wakeups (tests/kernel/context cpu_idle / timer_interrupts) and is a power regression versus the dlist and heap backends. Verified the timeout-functional suites (timer_api, timeout, timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86, x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected. The timer-wheel data structure, bucketing scheme and sift algorithm are the work of Peter Mitsis (PR #108339), re-homed here behind the timeout backend interface. Co-authored-by: Peter Mitsis <peter.mitsis@intel.com> Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-20 00:45:42 -04:00
#endif /* _TIMEOUT_BACKEND_OWNS_ANNOUNCE */
kernel: timeout: make in-announce check CPU-aware Several pieces of state in kernel/timeout.c are read by code paths on both the announcing CPU and other CPUs: announce_remaining, the "in-announce" lock that prevents parallel firing on SMP, and elapsed()/the +1 round-up in z_add_timeout(). Originally all of these were keyed off announce_remaining != 0, which conflates two distinct roles -- a remaining-ticks counter and a we-are-announcing flag -- and breaks down on SMP and across same-tick groups. Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads running on CPU B see announce_remaining != 0 and incorrectly take the in-announce branch -- elapsed() returns 0 so curr_tick + elapsed() goes backwards relative to a reading taken before the announce started, and a slice rearm or k_timer_start() coming from CPU B fires one tick early because the +1 round-up is suppressed even though CPU B is partway through a tick. The same condition was the root cause of the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and /a320 once the timer accounting cleanup in #107452 stopped masking it. Fix: distinguish "this CPU is at a tick edge (announcing)" from "some CPU is in the firing loop". Add announcing_cpu, set to the CPU id at the top of sys_clock_announce_locked() and cleared at the bottom; expose it via this_cpu_announcing() and any_cpu_announcing(). The former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s +1 round-up suppression -- both correct only on the announcing CPU. The latter drives the SMP early-return and z_add_timeout()'s deferred sys_clock_set_timeout call -- both correct for any CPU observing that an announce is in progress, and both robust to announce_remaining hitting 0 mid-loop (which now becomes possible -- see symptom 2 below). elapsed() on non-announcing CPUs now returns sys_clock_elapsed() + announce_remaining. The driver bumped its announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK at ISR entry while the kernel has only advanced curr_tick by the ticks already committed in the loop, so announce_remaining is the residual that must be added to sys_clock_elapsed() to recover T_real - curr_tick across the announce window. The new invariant is curr_tick + announce_remaining + sys_clock_elapsed() == T_real on every CPU at every point inside or outside the announce window. Symptom 2: that invariant relies on curr_tick and announce_remaining moving in lock-step. The original loop advanced curr_tick at the top of each outer iteration but only decremented announce_remaining after the handler returned, with the lock dropped around the handler. A non-announcing CPU calling sys_clock_tick_get() during a handler therefore observed curr_tick + announce_remaining one batch (dt) into the future, breaking the invariant for the duration of the handler. Fix: pair the updates so they happen under the same lock acquisition before the handler runs: curr_tick += t->dticks; announce_remaining -= t->dticks; Symptom 3: commit d157b3da193d ("kernel: timeout: keep announce_remaining stable across same-tick group") kept announce_remaining > 0 across the inner drain of a same-tick group so that (a) the SMP early-return would still detect "another CPU is announcing" and (b) elapsed() in same-tick handlers would still anchor to the firing tick. With the lock-state role moved to announcing_cpu, neither property depends on announce_remaining anymore: any_cpu_announcing() carries (a) and this_cpu_announcing() carries (b), both throughout the loop on the announcing CPU regardless of the budget. The inner drain do-while loop is therefore no longer needed -- the outer loop's t->dticks <= announce_remaining naturally accepts dticks == 0 follow-on timeouts even when announce_remaining has reached 0. Fixes #106317 Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-05 22:27:18 -04:00
announcing_cpu = -1;
reprogram_next(0);
k_spin_unlock(&timeout_lock, key);
kernel/sched: Use kernel timeouts for timeslice expirations Rework the fragile and ad-hoc computation of timeslice expirations into per-CPU struct _timeout objects with regular callbacks. The expiration callbacks themselves simply set a per-cpu flag (they might run on any CPU), which gets checked at the end of the timer ISR on every CPU. This simplifies logic and removes a bunch of code. It also fixes at least three bugs: 1. As @npitre discovered: On SMP, the number of ticks announced on any given CPU is going to be a subset of all expired ticks. This broke the accounting of timeslice ticks, and effectively meant that timeslicing only worked on SMP on systems where one CPU could hog all the announcements, and only on that CPU. 2. The bootstrap path to arm the timer driver after setting the first timeout in an empty list couldn't take into account sys_clock_elapsed() ticks, as it didn't know whether it was being called underneath an existing announce loop. Now this code is no longer responsible for knowing anything about time slicing at all. 3. Also on SMP, there was a case where two CPUs timeslicing simultaneously could stomp on each others' timeouts in z_set_timeout_expiry(), as neither had a way of knowing what the other's state was. CPUs could miss their own expiration and have to wait for the slice expiration on the other CPU. Now, timeouts are global objects with simple expiration times, and there's no need for that function at all. Signed-off-by: Andy Ross <andyross@google.com>
2023-03-06 14:31:35 -08:00
#ifdef CONFIG_TIMESLICING
z_time_slice();
#endif /* CONFIG_TIMESLICING */
}
kernel/timeout: introduce sys_clock_lock() and sys_clock_announce_locked() On SMP systems with tickless kernels, a race condition exists between timer driver ISRs and the kernel's tick accounting. The driver updates its hardware cycle baseline under a private lock, then calls sys_clock_announce() which updates curr_tick under the separate timeout_lock. In the gap between these two lock releases, any kernel code calling sys_clock_elapsed() sees the new driver baseline but the old curr_tick, producing inconsistent time values that can go backwards. This affects every code path using the internal elapsed() helper: uptime queries, timeout scheduling, timeout cancellation, remaining time queries, and next-expiry calculations. The root cause is two separate locks protecting state that must be mutually consistent. Fix this by exposing the kernel's timeout_lock to timer drivers via sys_clock_lock()/sys_clock_unlock(), and providing sys_clock_announce_locked() which assumes the lock is already held. Timer drivers can now acquire the single lock, update their hardware state, and announce ticks all under the same lock — eliminating the race window entirely. The key is passed to sys_clock_announce_locked() which consumes it (releasing the lock when it returns). The existing sys_clock_announce() becomes a backward-compatible wrapper, allowing incremental driver migration with no flag day. Document that sys_clock_set_timeout(), sys_clock_elapsed(), and sys_clock_idle_exit() are called by the kernel with the timer lock held. Update the timer driver guide in clocks.rst accordingly. Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-03-13 16:43:25 -04:00
#if defined(CONFIG_SMP) || defined(CONFIG_SPIN_VALIDATE)
k_spinlock_key_t sys_clock_lock(void)
{
return k_spin_lock(&timeout_lock);
}
void sys_clock_unlock(k_spinlock_key_t key)
{
k_spin_unlock(&timeout_lock, key);
}
#endif
#if defined(CONFIG_TEST) || defined(CONFIG_ASSERT)
bool sys_clock_is_locked(void)
{
return z_spin_is_locked(&timeout_lock);
}
#endif
int64_t sys_clock_tick_get(void)
{
uint64_t t = 0U;
K_SPINLOCK(&timeout_lock) {
t = curr_tick + elapsed();
}
return t;
}
uint32_t sys_clock_tick_get_32(void)
{
#ifdef CONFIG_TICKLESS_KERNEL
return (uint32_t)sys_clock_tick_get();
#else
return (uint32_t)curr_tick;
#endif /* CONFIG_TICKLESS_KERNEL */
}
int64_t z_impl_k_uptime_ticks(void)
{
return sys_clock_tick_get();
}
#ifdef CONFIG_USERSPACE
static inline int64_t z_vrfy_k_uptime_ticks(void)
{
return z_impl_k_uptime_ticks();
}
#include <zephyr/syscalls/k_uptime_ticks_mrsh.c>
#endif /* CONFIG_USERSPACE */
kernel/timeout: Make timeout arguments an opaque type Add a k_timeout_t type, and use it everywhere that kernel API functions were accepting a millisecond timeout argument. Instead of forcing milliseconds everywhere (which are often not integrally representable as system ticks), do the conversion to ticks at the point where the timeout is created. This avoids an extra unit conversion in some application code, and allows us to express the timeout in units other than milliseconds to achieve greater precision. The existing K_MSEC() et. al. macros now return initializers for a k_timeout_t. The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t values, which means they cannot be operated on as integers. Applications which have their own APIs that need to inspect these vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to test for equality. Timer drivers, which receive an integer tick count in ther z_clock_set_timeout() functions, now use the integer-valued K_TICKS_FOREVER constant instead of K_FOREVER. For the initial release, to preserve source compatibility, a CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the k_timeout_t will remain a compatible 32 bit value that will work with any legacy Zephyr application. Some subsystems present timeout (or timeout-like) values to their own users as APIs that would re-use the kernel's own constants and conventions. These will require some minor design work to adapt to the new scheme (in most cases just using k_timeout_t directly in their own API), and they have not been changed in this patch, instead selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems include: CAN Bus, the Microbit display driver, I2S, LoRa modem drivers, the UART Async API, Video hardware drivers, the console subsystem, and the network buffer abstraction. k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant provided that works identically to the original API. Most of the changes here are just type/configuration management and documentation, but there are logic changes in mempool, where a loop that used a timeout numerically has been reworked using a new z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was enabled) a similar loop was needlessly used to try to retry the k_poll() call after a spurious failure. But k_poll() does not fail spuriously, so the loop was removed. Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
k_timepoint_t sys_timepoint_calc(k_timeout_t timeout)
kernel/timeout: Make timeout arguments an opaque type Add a k_timeout_t type, and use it everywhere that kernel API functions were accepting a millisecond timeout argument. Instead of forcing milliseconds everywhere (which are often not integrally representable as system ticks), do the conversion to ticks at the point where the timeout is created. This avoids an extra unit conversion in some application code, and allows us to express the timeout in units other than milliseconds to achieve greater precision. The existing K_MSEC() et. al. macros now return initializers for a k_timeout_t. The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t values, which means they cannot be operated on as integers. Applications which have their own APIs that need to inspect these vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to test for equality. Timer drivers, which receive an integer tick count in ther z_clock_set_timeout() functions, now use the integer-valued K_TICKS_FOREVER constant instead of K_FOREVER. For the initial release, to preserve source compatibility, a CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the k_timeout_t will remain a compatible 32 bit value that will work with any legacy Zephyr application. Some subsystems present timeout (or timeout-like) values to their own users as APIs that would re-use the kernel's own constants and conventions. These will require some minor design work to adapt to the new scheme (in most cases just using k_timeout_t directly in their own API), and they have not been changed in this patch, instead selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems include: CAN Bus, the Microbit display driver, I2S, LoRa modem drivers, the UART Async API, Video hardware drivers, the console subsystem, and the network buffer abstraction. k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant provided that works identically to the original API. Most of the changes here are just type/configuration management and documentation, but there are logic changes in mempool, where a loop that used a timeout numerically has been reworked using a new z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was enabled) a similar loop was needlessly used to try to retry the k_poll() call after a spurious failure. But k_poll() does not fail spuriously, so the loop was removed. Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
{
k_timepoint_t timepoint;
kernel/timeout: Make timeout arguments an opaque type Add a k_timeout_t type, and use it everywhere that kernel API functions were accepting a millisecond timeout argument. Instead of forcing milliseconds everywhere (which are often not integrally representable as system ticks), do the conversion to ticks at the point where the timeout is created. This avoids an extra unit conversion in some application code, and allows us to express the timeout in units other than milliseconds to achieve greater precision. The existing K_MSEC() et. al. macros now return initializers for a k_timeout_t. The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t values, which means they cannot be operated on as integers. Applications which have their own APIs that need to inspect these vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to test for equality. Timer drivers, which receive an integer tick count in ther z_clock_set_timeout() functions, now use the integer-valued K_TICKS_FOREVER constant instead of K_FOREVER. For the initial release, to preserve source compatibility, a CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the k_timeout_t will remain a compatible 32 bit value that will work with any legacy Zephyr application. Some subsystems present timeout (or timeout-like) values to their own users as APIs that would re-use the kernel's own constants and conventions. These will require some minor design work to adapt to the new scheme (in most cases just using k_timeout_t directly in their own API), and they have not been changed in this patch, instead selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems include: CAN Bus, the Microbit display driver, I2S, LoRa modem drivers, the UART Async API, Video hardware drivers, the console subsystem, and the network buffer abstraction. k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant provided that works identically to the original API. Most of the changes here are just type/configuration management and documentation, but there are logic changes in mempool, where a loop that used a timeout numerically has been reworked using a new z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was enabled) a similar loop was needlessly used to try to retry the k_poll() call after a spurious failure. But k_poll() does not fail spuriously, so the loop was removed. Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
if (K_TIMEOUT_EQ(timeout, K_FOREVER)) {
timepoint.tick = UINT64_MAX;
kernel/timeout: Make timeout arguments an opaque type Add a k_timeout_t type, and use it everywhere that kernel API functions were accepting a millisecond timeout argument. Instead of forcing milliseconds everywhere (which are often not integrally representable as system ticks), do the conversion to ticks at the point where the timeout is created. This avoids an extra unit conversion in some application code, and allows us to express the timeout in units other than milliseconds to achieve greater precision. The existing K_MSEC() et. al. macros now return initializers for a k_timeout_t. The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t values, which means they cannot be operated on as integers. Applications which have their own APIs that need to inspect these vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to test for equality. Timer drivers, which receive an integer tick count in ther z_clock_set_timeout() functions, now use the integer-valued K_TICKS_FOREVER constant instead of K_FOREVER. For the initial release, to preserve source compatibility, a CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the k_timeout_t will remain a compatible 32 bit value that will work with any legacy Zephyr application. Some subsystems present timeout (or timeout-like) values to their own users as APIs that would re-use the kernel's own constants and conventions. These will require some minor design work to adapt to the new scheme (in most cases just using k_timeout_t directly in their own API), and they have not been changed in this patch, instead selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems include: CAN Bus, the Microbit display driver, I2S, LoRa modem drivers, the UART Async API, Video hardware drivers, the console subsystem, and the network buffer abstraction. k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant provided that works identically to the original API. Most of the changes here are just type/configuration management and documentation, but there are logic changes in mempool, where a loop that used a timeout numerically has been reworked using a new z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was enabled) a similar loop was needlessly used to try to retry the k_poll() call after a spurious failure. But k_poll() does not fail spuriously, so the loop was removed. Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
} else if (K_TIMEOUT_EQ(timeout, K_NO_WAIT)) {
timepoint.tick = 0;
} else {
k_ticks_t dt = timeout.ticks;
if (Z_IS_TIMEOUT_RELATIVE(timeout)) {
timepoint.tick = sys_clock_tick_get() + max(1, dt);
} else {
timepoint.tick = Z_TICK_ABS(dt);
}
}
return timepoint;
}
k_timeout_t sys_timepoint_timeout(k_timepoint_t timepoint)
{
uint64_t now, remaining;
if (timepoint.tick == UINT64_MAX) {
return K_FOREVER;
}
if (timepoint.tick == 0) {
return K_NO_WAIT;
}
now = sys_clock_tick_get();
remaining = (timepoint.tick > now) ? (timepoint.tick - now) : 0;
return K_TICKS(remaining);
kernel/timeout: Make timeout arguments an opaque type Add a k_timeout_t type, and use it everywhere that kernel API functions were accepting a millisecond timeout argument. Instead of forcing milliseconds everywhere (which are often not integrally representable as system ticks), do the conversion to ticks at the point where the timeout is created. This avoids an extra unit conversion in some application code, and allows us to express the timeout in units other than milliseconds to achieve greater precision. The existing K_MSEC() et. al. macros now return initializers for a k_timeout_t. The K_NO_WAIT and K_FOREVER constants have now become k_timeout_t values, which means they cannot be operated on as integers. Applications which have their own APIs that need to inspect these vs. user-provided timeouts can now use a K_TIMEOUT_EQ() predicate to test for equality. Timer drivers, which receive an integer tick count in ther z_clock_set_timeout() functions, now use the integer-valued K_TICKS_FOREVER constant instead of K_FOREVER. For the initial release, to preserve source compatibility, a CONFIG_LEGACY_TIMEOUT_API kconfig is provided. When true, the k_timeout_t will remain a compatible 32 bit value that will work with any legacy Zephyr application. Some subsystems present timeout (or timeout-like) values to their own users as APIs that would re-use the kernel's own constants and conventions. These will require some minor design work to adapt to the new scheme (in most cases just using k_timeout_t directly in their own API), and they have not been changed in this patch, instead selecting CONFIG_LEGACY_TIMEOUT_API via kconfig. These subsystems include: CAN Bus, the Microbit display driver, I2S, LoRa modem drivers, the UART Async API, Video hardware drivers, the console subsystem, and the network buffer abstraction. k_sleep() now takes a k_timeout_t argument, with a k_msleep() variant provided that works identically to the original API. Most of the changes here are just type/configuration management and documentation, but there are logic changes in mempool, where a loop that used a timeout numerically has been reworked using a new z_timeout_end_calc() predicate. Also in queue.c, a (when POLL was enabled) a similar loop was needlessly used to try to retry the k_poll() call after a spurious failure. But k_poll() does not fail spuriously, so the loop was removed. Signed-off-by: Andy Ross <andrew.j.ross@intel.com>
2020-03-05 15:18:14 -08:00
}
#ifdef CONFIG_ZTEST
void z_impl_sys_clock_tick_set(uint64_t tick)
{
curr_tick = tick;
}
void z_vrfy_sys_clock_tick_set(uint64_t tick)
{
z_impl_sys_clock_tick_set(tick);
}
#endif /* CONFIG_ZTEST */