Commit graph zephyr/kernel
Author SHA1 Message Date
Nicolas Pitre
ce23f1d734 kernel: sched: switch z_thread_timeout() to the superseded check
With every caller of z_abort_thread_timeout() now migrated to
z_try_abort_thread_timeout() (sched.c, thread.c, scheduler.c,
events.c, pipe.c) or to z_unpend_first_thread_locked() (sem.c,
mutex.c, mem_slab.c, stack.c, condvar.c, msg_q.c, queue.c, futex.c),
remove the inline wrapper (and !CONFIG_SYS_CLOCK_EXISTS stub) for
z_abort_thread_timeout().

z_thread_timeout() still needs a cancellation check, but it no longer
relies on the dticks=ANNOUNCING sentinel: switch it to
z_timeout_inflight_superseded(). The check is still required, and for
the same reason 1b8c7a3 added it. A concurrent waker on another CPU
(e.g. a sem give via z_unpend_first_thread_locked()) can unpend and
ready the thread while this timeout's handler is blocked on
_sched_spinlock; the thread may then run and re-pend on a different
object -- possibly with no timeout (K_FOREVER). The waker aborts this
timeout, which flags it superseded, and z_thread_timeout() bails on
that flag so it does not wake the thread from its new wait. The
atomic wake-under-_sched_spinlock closes the swap_retval window; the
superseded check closes this re-pend window.

The remaining TIMEOUT_DTICKS_ANNOUNCING sentinel and
z_is_timeout_handler_canceled() helper still have other users
(kernel/timer.c, kernel/poll.c, kernel/work.c) and are removed in the
later cleanup commit once those subsystems have been migrated as well.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
cf6d9d8fa8 kernel: wait_q: add z_unpend_first_thread_locked() and migrate callers
Replace z_unpend_first_thread() with z_unpend_first_thread_locked() and
migrate every caller across the kernel. The old function dropped the
scheduler spinlock before returning, exposing a race window between
its caller's "arch_thread_return_value_set + z_ready_thread" pair and
a still-in-flight timeout handler that could ready the thread first --
the woken thread might then run on another CPU and see an uninitialized
swap_retval. Pre-1b8c7a3 the dticks-cancel check made the handler bail;
here we fix it cleanly by requiring the caller to hold _sched_spinlock
across the entire wake, so the handler is blocked for the duration and
runs as a no-op afterwards.

z_unpend_first_thread_locked() requires the caller to be inside a
locked region and must be paired with z_sched_ready_locked() (and
whatever return-value setup is needed) under the same lock acquisition.

Sites migrated:

  Simple "set retval [+ swap_data] and ready" callers use the existing
  z_sched_wake() convenience wrapper, refactored to use the new
  helper internally:
    sem (give, reset), mem_slab (free), stack (push),
    condvar (signal, broadcast), msgq (purge),
    queue (cancel_wait, queue_insert, append_list),
    futex (wake).

  Sites that need additional setup on the woken thread use
  LOCK_SCHED_SPINLOCK + z_unpend_first_thread_locked() + custom wake:
    mutex (unlock -- needs the thread reference to track new owner),
    msgq put / get (needs memcpy into the receiver's swap_data
                    buffer before the return value is set).

The dticks-cancel check in z_thread_timeout() is left in place; it is
no longer load-bearing once z_abort_thread_timeout() has no callers,
and is removed in the next commit.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
387b387479 kernel: sched: migrate scheduler-internal sites to z_try_abort_timeout()
Migrate scheduler-internal callers of z_abort_thread_timeout() to the
new z_try_abort_thread_timeout(). This covers the abort sites in
sched.c, thread.c, scheduler.c (z_sched_wake), events.c, and pipe.c.

The patterns used:

  z_unpend_thread (sched.c) retries on -EAGAIN: if the timeout
  handler is in flight on another CPU, drop _sched_spinlock so the
  handler can run to completion and retry. This preserves 1b8c7a3's
  unpend+abort atomicity from the caller's perspective.

  halt_thread (sched.c) takes the caller's sched-lock key as a pointer
  so its direct abort on the dying thread can retry on -EAGAIN.
  Waiting for the handler is mandatory: a caller may free the thread's
  storage as soon as halt_thread() returns, and without waiting, the
  still-in-flight handler would later dereference freed memory.
  _THREAD_DEAD is set before the abort, so the handler bails via the
  killed check in z_sched_wake_thread_locked().

  For next_up() (the scheduler-context caller), the key is not cleanly
  available: K_SPINLOCK in z_get_next_switch_handle and do_swap's
  (void)k_spin_lock both discard it. halt_thread is invoked on
  _current with NULL key and no abort is performed -- _current is
  running, so its base.timeout cannot be linked. The gap is closed at
  the other end: z_thread_halt() spins on z_try_abort_thread_timeout()
  outside any lock after the halt-queue wait completes, before
  returning to the caller of k_thread_abort().

  z_unpend_all_locked / unpend_all (sched.c) skip the local
  ready_thread() on -EAGAIN and let the still-blocked handler ready
  the thread when _sched_spinlock drops. The threads being woken are
  not freed, so no UAF risk; end state is identical.

  z_impl_k_wakeup (thread.c), z_sched_wake (scheduler.c) and
  event_walk_op (events.c) perform the wake entirely under
  _sched_spinlock, so a (void) abort is race-free -- a racing in-flight
  handler is blocked on the same lock during the wake.

  copy_to_pending_readers (pipe.c) is restructured to also wake the
  reader under the scheduler lock instead of after it, so the
  return-value set, unpend, abort, and ready all happen atomically.

The dticks-cancel check in z_thread_timeout() is preserved for now
because other callers (sem.c, mutex.c, ... via z_unpend_first_thread())
still use z_abort_thread_timeout() and rely on it for race protection.
A follow-up commit migrates those, and a final commit drops the
cancel check and removes z_abort_thread_timeout() itself.

Add z_try_abort_thread_timeout() as an inline wrapper.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
6d83c1abd0 kernel: timer: migrate to z_try_abort_timeout()
Switch z_impl_k_timer_start() and z_impl_k_timer_stop() to the new
z_try_abort_timeout() interface.

Neither waits for an in-flight handler. When z_try_abort_timeout()
returns non-zero (the timeout was not on the queue), the handler is
either inactive or dispatching on another CPU; in the latter case it
has been flagged superseded and will bail at its entry check. This
matches the best-effort behaviour of the legacy z_abort_timeout():
on main, a cross-CPU stop set dticks=ABORTED and returned without
waiting, and the handler bailed on the sentinel. The superseded bit
now carries that signal.

Not waiting is also required for correctness: a k_timer expiry_fn
runs arbitrary user code, which may block on the very CPU that is
trying to stop the timer (e.g. k_thread_abort() of a thread running
there). A mandatory wait-for-handler in k_timer_stop()/start() would
deadlock that case; the only caller that must wait is k_timer_cleanup()
(about to free the storage), which keeps its retry loop.

k_timer_start: the lock encapsulation around abort + add is preserved
(it serializes concurrent k_timer_start on the same timer). The re-arm
re-links the node, which makes a not-yet-committed handler bail via
the sys_dnode_is_linked check too.

k_timer_stop: a non-zero return means "nothing on the queue to stop",
so stop_fn and the wait_q wake run only when the timeout was actually
dequeued (return 0).

z_timer_expiration_handler() drops the z_is_timeout_handler_canceled
check and bails on sys_dnode_is_linked (a higher-priority interrupt
between sys_clock_announce()'s unlock and the handler taking
timer.c::lock can re-link the timeout via k_timer_start) or on
z_timeout_inflight_superseded() (the timeout was aborted while
dispatching). Skipping avoids a double expiry and the periodic
restart asserting in z_add_timeout() on an already-linked node.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
61444b281e kernel: timeout: tag inflight_timeout with a superseded bit
When z_try_abort_timeout() finds the target timeout already popped
from the queue and in flight (its handler dispatching), the abort
cannot remove anything from the queue. The handler is about to run
(or is running) the timeout's callback; the aborter needs a way to
tell it "you were aborted, skip your side effects". main carried that
signal in the per-timeout dticks=ABORTED sentinel, which any CPU
could write under timeout_lock and the handler checked at entry via
z_is_timeout_handler_canceled(). Later commits in this series remove
the dticks-cancel mechanism, so the signal needs a new home.

Encode it in the low bit of the file-local inflight_timeout pointer
(struct _timeout is pointer-aligned, so bit 0 is free):

    inflight_timeout == NULL       no handler in flight
    inflight_timeout == t          handler in flight, not superseded
    inflight_timeout == t | 1      handler in flight, superseded

z_try_abort_timeout() sets the bit whenever the target is the
in-flight timeout -- on the announcing CPU (same-CPU IRQ that
preempted the dispatch, or a stop after a re-arm) and on another CPU
racing the handler.

A handler with non-idempotent side effects (currently only k_timer's
expiry_fn) checks z_timeout_inflight_superseded() at entry and bails
if set. Idempotent handlers (z_thread_timeout, work, poll, ...)
tolerate the race and don't need the check.

The bit is a best-effort signal, not a barrier: an aborter that sets
it after the handler has passed its check has no effect (the handler
already committed), exactly as the dticks sentinel behaved on main.
The -EAGAIN return for the cross-CPU case is a separate mechanism,
used only by callers that must wait for the handler to fully complete
(e.g. before freeing the timeout's storage); they spin, while
best-effort callers ignore -EAGAIN and rely on the bit.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
805b2af18e kernel: timeout: introduce z_try_abort_timeout() with in-flight tracking
Add an `inflight_timeout` pointer in the timeout subsystem, set by
sys_clock_announce_locked() under timeout_lock before a handler is
dispatched and cleared after the handler returns. The single-announcer
SMP invariant (only one CPU is in the dispatch loop at a time) means a
single global pointer suffices, mirroring announcing_cpu.

Add a new abort function z_try_abort_timeout() that uses this pointer
to detect "popped from the queue but handler not yet finished" without
relying on the in-band dticks == ANNOUNCING sentinel that survives
across the timeout's storage being freed only by coincidence.

The new function returns -EAGAIN when a handler is in flight on
another CPU, signalling that the caller must drop any outer lock the
handler may need and retry. A short arch_spin_relax() is performed
before -EAGAIN is returned so callers don't need to add their own
back-off in the retry path. inflight_timeout itself is kept private
to the timeout subsystem; callers only ever see the return value.

This commit introduces the infrastructure but migrates no callers.
The existing z_abort_timeout() and TIMEOUT_DTICKS_ANNOUNCING-based
cancellation in handlers continue to work; the dispatch loop sets
both inflight_timeout and ANNOUNCING so the two paths coexist while
callers are migrated one module at a time.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Jinming Zhao
d844a4a861 kernel: spinlock: add z_assert_can_swap() for context-switch validation
Replace the single in-place assertion in do_swap/z_swap_irqlock with a
z_assert_can_swap() helper performing three checks: held spinlock
pointer, hold count, and IRQ state.

To support the hold-count check, add per-CPU tracking arrays
(z_held_spinlock[] and z_held_spinlock_count[]).

Add z_spin_lock_transfer_owner() to update lock ownership after a
context switch, z_spinlock_abort_sentinel to exempt threads aborted by
a ztest expected-fault scenario, and z_spin_validate_reset() to reset
stale per-CPU lock tracking left by such an abort so subsequent tests
can proceed cleanly.

Signed-off-by: Jinming Zhao <jinmzhao@qti.qualcomm.com>
2026-06-15 18:35:34 -04:00
Anas Nashif
9f788430c9 Revert "kernel: mutex: detect re-init and uninitialized use via magic sentinel"
This reverts commit ff508efd6a.

This is causing breakages across the tree. We should fix all issues and
retry. Nothing wrong with the change itself, but the tree was not
prepared for this change.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-10 14:54:00 -04:00
James Roy
465d3de809 kernel: Add k_thread_runtime_stats_is_enabled function
Add 'k_thread_runtime_stats_is_enabled' function, whichs
used to check whether runtime statistics collection is
enabled for a thread.

Signed-off-by: James Roy <rruuaanng@outlook.com>
2026-06-10 14:52:53 -04:00
Mayur Salve
ff508efd6a kernel: mutex: detect re-init and uninitialized use via magic sentinel
Add a uintptr_t magic field (K_MUTEX_MAGIC = K_OBJ_TYPE_MUTEX_ID) to
struct k_mutex under CONFIG_ASSERT, written by k_mutex_init() and
Z_MUTEX_INITIALIZER. Assert the sentinel in lock/unlock to catch
use-before-init, and assert the mutex is not held before re-init.

Zero-initialize dynamically allocated objects in dynamic_object_create()
under CONFIG_ASSERT for a known starting state. Fix test bugs where
k_mutex_init() was called on a held mutex. Add assertion path tests.

Signed-off-by: Mayur Salve <msalve@qti.qualcomm.com>
2026-06-10 13:11:47 +02:00
Anas Nashif
7798570a03 kernel: sched: enforce non-zero CPU mask invariant in PIN_ONLY mode
In CONFIG_SCHED_CPU_MASK_PIN_ONLY a thread must always be pinned to
exactly one CPU.  Two related gaps let a zero-mask thread slip
through silently:

1. thread_runq() (run_q.h) had an explicit if/else that silently
   routed a zero-masked thread to CPU 0 instead of catching the
   violation.  Replace it with an __ASSERT that fires at the point
   where the bad state is used.

2. cpu_mask_mod() (cpu_mask.c) checked '(m == 0) || power-of-two',
   which accepted a cleared mask as valid.  Tighten the check to
   require strictly a single set bit ('m != 0 && power-of-two') so
   any API call that would leave the thread with an empty mask (e.g.
   k_thread_cpu_mask_clear()) traps at the API boundary rather than
   later at queue time.

The two enforcement points now agree: every thread in PIN_ONLY mode
must carry exactly one CPU bit from the moment the mask is written
until the thread is queued.

Assisted-by: GitHub Copilot:claude-sonnet-4-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-09 19:57:09 +02:00
Anas Nashif
50f4e51efb kernel: smp: update SCHED_CPU_MASK for multi-scheduler support
Now SCHED_CPU_MASK is no longer restricted to SCHED_SIMPLE. Dropped
dependency on SCHED_SIMPLE and made SCHED_CPU_MASK work on all scheduler
types by implementing z_priq_rb_mask_best and z_priq_mq_mask_best for
both scalable and multiq schedulers.

Update the help text and remove the stale claim that the feature only
works with the simple scheduler and document the mask-aware best-thread
search algorithm and its performance characteristics for each of the
three supported backends:

  SCHED_SIMPLE:   O(N) sorted-list scan; terminates at the first
                  priority-eligible thread, so cost is proportional to
                  the number of higher-priority threads pinned away from
                  the current CPU.

  SCHED_SCALABLE: O(N) in-order rbtree walk; benefits from priority
                  ordering so the walk usually terminates early.

  SCHED_MULTIQ:   O(P) bitmap iteration over non-empty priority levels
                  plus an inner O(N) per-level list scan; fast when
                  affinity-constrained threads are sparse, degrades when
                  many same-priority threads are pinned away.

Assisted-by: GitHub Copilot:claude-sonnet-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-09 19:57:09 +02:00
Peter Mitsis
6e9d4b7f33 kernel: Fix a race in k_queue_unique_append()
Fixes a TOCTOU race in k_queue_unique_append() by locking the
queue's spinlock around both the search and insert operations.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-08 23:13:31 +02:00
Peter Mitsis
bb870f2e8f kernel: obtain lock before queue_insert()
The k_queue helper routine queue_insert() now requires that
the queue's spinlock be obtained prior to calling it. Its
key is now passed as a parameter into it so that the lock
will be released before it returns.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-08 23:13:31 +02:00
Peter Mitsis
3fc6d9e700 kernel: k_queue_remove() locks spinlock
To avoid a TOCTOU type error in k_queue_remove(), it must lock
the queue's spinlock for the duration of the operation.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-08 23:13:31 +02:00
Peter Mitsis
a6b6149a50 kernel: z_queue_node_peek() needs spinlock held
When peeking at an allocated node, the allocated node must be
dereferenced. Unless the queue's spinlock is held that allocated
node could be freed and that memory re-used for something else
entirely leading to the kernel dereferencing an invalid pointer.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-08 23:13:31 +02:00
Peter Mitsis
6ec1b773f1 kernel: Make z_queue_node_peek() static
The routine z_queue_node_peek() is only referenced from within
queue.c. Make it static.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-08 23:13:31 +02:00
Fabio Baltieri
63732a6fb5 device: make DEVICE_API_GET assert optionals
Asserts on DEVICE_API_GET can take a significant amount of flash, to the
point of making an imagine not fit the flash anymore, add an option to
disable them.

Signed-off-by: Fabio Baltieri <fabiobaltieri@google.com>
2026-06-04 14:05:26 +02:00
Benjamin Cabé
5198d1ea87 doc: kernel: add Doxygen stubs for ARCH_DATA_PAGE_* flags
Add __DOXYGEN__ placeholder definitions for ARCH_DATA_PAGE_* macros so
that Doxygen, which does not know about how these macros _might_ be
defined by architectures, can still bind documentation to them.

Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
2026-06-03 18:28:13 -04:00
Sofian Elmotiem
f77f55fb16 kernel/pipe: fix swap_data corruption when k_pipe_read is called from ISR
In ISR context _current is the interrupted thread, not the ISR itself.
Setting _current->base.swap_data from an ISR corrupts a field that
belongs to that thread and may be in active use.

ISR callers must use K_NO_WAIT and never pend, so they never need the
direct-copy buffer. Moving the swap_data assignment into wait_for()
after the K_NO_WAIT early-return ensures it is only set on the path
that will actually pend, which is never the ISR path.

Fixes: #110077

Signed-off-by: Sofian Elmotiem <sofianelmotiem@gmail.com>
2026-06-02 20:25:34 +02:00
Anas Nashif
d0b389da9e kernel: remove redundant kernel_structs.h includes
kernel.h implies kernel_structs.h via kernel_includes.h, making
explicit inclusion of kernel_structs.h unnecessary whenever kernel.h
is already included in the same translation unit.

Remove the redundant includes across arch, boards, drivers, kernel,
lib, samples, subsys, and tests trees.

in include/zephyr/kernel_structs.h:
 *  2. kernel.h shall imply kernel_structs.h, such that it shall not be
 *    necessary to include kernel_structs.h explicitly when kernel.h is
 *    included.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-02 20:24:14 +02:00
Mugunthan V
2ba2fa6f9b kernel: sched: enforce half-modulus invariant in k_thread_deadline_set()
The EDF scheduler comparator z_sched_prio_cmp() relies on a half-modulus
invariant where all active absolute deadlines must lie within the same
2^31-cycle window to ensure correct signed comparison.

Clamping relative deadlines to INT_MAX violated this. For example, two
threads setting deadlines 1 cycle apart near 0x80000000 could produce
absolute deadlines at 0x80000000 and 0x00000000, causing a priority
inversion in the comparator (evaluating d2 - d1 to -2^31).

Fix this by:
1. Tightening the relative deadline clamp to INT32_MAX / 2 (2^30 cycles).
   This preserves the modular comparison and leaves 2^30 cycles of skew
   tolerance for active threads in the queue.
2. Documenting this cap in the public API Doxygen in kernel.h.
3. Cleaning up the addition to use defined uint32_t arithmetic.

Note: z_impl_k_thread_absolute_deadline_set() has no clamp and remains
an out-of-scope gap acknowledged here.

Signed-off-by: Mugunthan V <mugunthan@aerlync.com>
2026-06-02 20:21:42 +02:00
Sylvio Alves
c4b50e3bf3 kernel: smp: fix stale SCHED_CPU_MASK consumers
SCHED_CPU_MASK is an SMP-only feature. Several samples and
tests set CONFIG_SCHED_CPU_MASK=y unconditionally in prj.conf,
which warns on UP boards.

Move the assignment into sample.yaml or testcase.yaml gated on
CONFIG_SMP. Pair with CONFIG_SCHED_SIMPLE=y to satisfy the
scheduler dependency. Drop the flag from tests that do not use
the API. Remove a stale help-text paragraph that no longer
matches the symbol's placement.

Signed-off-by: Sylvio Alves <sylvio.alves@espressif.com>
2026-06-01 18:06:04 -05:00
Peter Mitsis
9f6de47ed2 syscall: Add K_OBJ_DRIVER_ANY
Adds a new generic kernel object type--K_OBJ_DRIVER_ANY. This is used
to validate that the specified object is a known driver type. Being
more restrictive than K_OBJ_ANY, it allows for better error checking
in both z_vrfy_device_init() and z_vrfy_device_is_ready().

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-05-29 20:50:29 -07:00
Nicolas Pitre
9912d8c3a5 kernel: timeout: micro-optimize elapsed() on UP
The "+ announce_remaining" added in commit 2ec65238d6 ("kernel:
timeout: make in-announce check CPU-aware") forced elapsed() to push
an 8-byte stack frame, as the compiler can no longer tail-call into
sys_clock_elapsed(). On UP, announce_remaining is always 0 when
this_cpu_announcing() is false, so the add is gratuitous there.

Gate it on CONFIG_SMP: with CONFIG_SMP=n the compiler folds the
addition away and restores the tail call.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-27 21:35:31 -04:00
Flavio Ceolin
862ea2fbbe kernel: userspace: fix thread_idx_alloc() for SMP
Hold lists_lock across the index allocation and permission clearing
to prevent races on SMP systems. For additional context:

https://github.com/zephyrproject-rtos/zephyr/pull/108721

Signed-off-by: Flavio Ceolin <flavio.ceolin@gmail.com>
2026-05-26 09:22:03 +02:00
Lingutla Chandrasekhar
a5fbcbb12c arch: riscv: skip ATOMIC_OPERATIONS_C if arch-specific atomics are enabled
Previously, ATOMIC_OPERATIONS_C was selected for RISC-V whenever the
'A' (atomic) ISA extension (RISCV_ISA_EXT_A) was absent. This caused
a conflict on platforms that lack the 'A' extension but still provide
their own arch-level atomic implementation via ATOMIC_OPERATIONS_ARCH
(e.g. future RISC-V SoCs with custom atomic support).

Add !ATOMIC_OPERATIONS_ARCH to the select condition so that the
generic C fallback (interrupt-locking) is only chosen when neither
the ISA extension nor an arch-specific implementation is available.

This condition creates a Kconfig dependency cycle:

  RISCV selects ATOMIC_OPERATIONS_C if !ATOMIC_OPERATIONS_ARCH
  => ATOMIC_OPERATIONS_C depends on !ATOMIC_OPERATIONS_ARCH
  => ATOMIC_OPERATIONS_ARCH depends on SMP (fvp_base_revc_2xaem board)
  => SMP depends on !ATOMIC_OPERATIONS_C

Break the cycle by removing 'depends on !ATOMIC_OPERATIONS_C' from
SMP in kernel/smp/Kconfig. This is safe because ATOMIC_OPERATIONS_C
is now only selected when ATOMIC_OPERATIONS_ARCH is absent, so the
two symbols are mutually exclusive by construction. The existing
BUILD_ASSERT(!IS_ENABLED(CONFIG_SMP)) in lib/os/atomic_c.c provides
a compile-time backstop against any misconfiguration.

Suggested-by: Nicolas Pitre <npitre@baylibre.com>
Signed-off-by: Lingutla Chandrasekhar <lingutla@qti.qualcomm.com>
2026-05-22 10:44:26 +02:00
Holt Sun
15af50d202 pm: keep irq restore ownership in idle
Keep the original architecture IRQ key owned by idle across a

successful system PM transition.

Add architecture hooks and the PM_STATE_SET_IRQ_LOCKED migration

contract for SoCs that keep PM hooks from unmasking interrupts.

Signed-off-by: Holt Sun <holt.sun@nxp.com>
2026-05-21 17:02:03 -04:00
Srikanth Patchava
5c6c6837cc kernel: fix k_condvar_wait mutex re-acquisition on timeout
Always re-acquire the mutex before returning from k_condvar_wait(),
even when z_pend_curr() returns a non-zero status (timeout or error).
The previous code only re-locked the mutex on success (ret == 0),
violating POSIX semantics which require the mutex to always be held
by the calling thread when the function returns.

The K_NO_WAIT early-return path now also keeps the mutex locked
(returning -EAGAIN immediately without touching the mutex), preserving
the calling-thread-holds-mutex invariant on every exit path.

Signed-off-by: Srikanth Patchava <srikanth.patchava@outlook.com>
2026-05-21 17:00:29 -04:00
Peter Mitsis
8dc7a37bc7 kernel: poll: z_vrfy_k_poll() to free memory
Reworks the K_OOPS(K_SYSCALL_OBJ(...)) logic in z_vrfy_k_poll()
to ensure that the allocated 'events_copy' is freed before the
K_OOPS() is performed.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-05-20 20:08:10 -04:00
Daniel Leung
04c58a2ddb kernel: userspace: fix validate_kernel_object type/init
When z_object_validate() is renamed to k_object_validate(),
the call to it inside validate_kernel_object() was incorrectly
modified, and reverted back using arguments on old version.
So fix that.

Signed-off-by: Daniel Leung <daniel.leung@intel.com>
2026-05-20 20:07:59 -04:00
Alberto Escolar Piedras
0b0b41ce9d kernel: Initialize timeout.dticks on k_timer_init()
Unlike Z_TIMER_INITIALIZER, k_timer_init() does not fully initialize the
provided timer.
This results on valgrind warning about a Conditional jump on uninitialised
value when calling k_timer_start() on that object later when dticks is
checked.

Let's initialize it to avoid this warning.

Signed-off-by: Alberto Escolar Piedras <alberto.escolar.piedras@nordicsemi.no>
2026-05-20 14:12:52 +02:00
Nicolas Pitre
c3f2a6a07a kernel: timeslicing: rearm with slice_size-1 when slicer just fired
z_time_slice_reset() armed the slicer with K_TICKS(slice_size), expecting
z_add_timeout()'s "+1" round-up to be skipped on tick-edge entry by the
this_cpu_announcing() check introduced in commit 2e2202af61 ("kernel:
timeout: make z_add_timeout round-up conditional on announce"). In
practice that check is false at every reachable z_time_slice_reset()
call site:

  - z_time_slice() runs from sys_clock_announce_locked()'s post-loop
    epilogue, after announce_remaining has been zeroed and the timeout
    lock has been released. The locking discipline doesn't permit
    rearming earlier (i.e. from inside the firing loop), so the
    timeout-edge property is lost by the time the slicer is rearmed.

  - On SMP, the slicer can fire on another CPU and IPI ours to do the
    actual scheduler work; the receiving CPU is even further from the
    announce window when it rearms.

  - update_cache() (non-SMP) and z_get_next_switch_handle() (SMP) both
    invoke z_time_slice_reset() from the dispatch path, also outside
    any announce.

The result is that every slicer fire produces a slice that's one tick
longer than configured -- k_sched_time_slice_set(N ms) ends up firing at
roughly N+tick ms.

Compensate explicitly: when slice_expired[cpu] is set we know the slicer
just fired and the new arm is conceptually tick-aligned, so pass
slice_size-1. The +1 round-up in z_add_timeout() then lands the next
fire at exactly slice_size ticks. Other reset paths (voluntary yield,
higher-prio preempt, thread creation) leave slice_expired clear and keep
the full "at least N ticks" rounding.

Also drop the redundant rearm in z_time_slice() when curr is being
swapped out: the dispatch path's z_time_slice_reset(new_thread) will
arm correctly with slice_expired still set, so the in-z_time_slice()
arm just installed a timeout that the dispatch path immediately
aborted and replaced. Keep it only when curr stays at the front of the
queue (single runnable thread at this priority), where the dispatch
path won't run.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-20 10:55:56 +02:00
Nicolas Pitre
2ec65238d6 kernel: timeout: make in-announce check CPU-aware
Several pieces of state in kernel/timeout.c are read by code paths on
both the announcing CPU and other CPUs: announce_remaining, the
"in-announce" lock that prevents parallel firing on SMP, and
elapsed()/the +1 round-up in z_add_timeout(). Originally all of these
were keyed off announce_remaining != 0, which conflates two distinct
roles -- a remaining-ticks counter and a we-are-announcing flag --
and breaks down on SMP and across same-tick groups.

Symptom 1 (issue #106317): on SMP, while CPU A is announcing, threads
running on CPU B see announce_remaining != 0 and incorrectly take the
in-announce branch -- elapsed() returns 0 so curr_tick + elapsed()
goes backwards relative to a reading taken before the announce
started, and a slice rearm or k_timer_start() coming from CPU B fires
one tick early because the +1 round-up is suppressed even though CPU B
is partway through a tick. The same condition was the root cause of
the test_slice_reset failure on fvp_base_revc_2xaem/v8a/smp/ns and
/a320 once the timer accounting cleanup in #107452 stopped masking it.

Fix: distinguish "this CPU is at a tick edge (announcing)" from "some
CPU is in the firing loop". Add announcing_cpu, set to the CPU id at
the top of sys_clock_announce_locked() and cleared at the bottom;
expose it via this_cpu_announcing() and any_cpu_announcing(). The
former drives elapsed()'s return-0 short-circuit and z_add_timeout()'s
+1 round-up suppression -- both correct only on the announcing CPU.
The latter drives the SMP early-return and z_add_timeout()'s
deferred sys_clock_set_timeout call -- both correct for any CPU
observing that an announce is in progress, and both robust to
announce_remaining hitting 0 mid-loop (which now becomes possible --
see symptom 2 below).

elapsed() on non-announcing CPUs now returns
sys_clock_elapsed() + announce_remaining. The driver bumped its
announced-cycles baseline to (curr_tick_initial + N) * CYC_PER_TICK
at ISR entry while the kernel has only advanced curr_tick by the
ticks already committed in the loop, so announce_remaining is the
residual that must be added to sys_clock_elapsed() to recover
T_real - curr_tick across the announce window. The new invariant is

    curr_tick + announce_remaining + sys_clock_elapsed() == T_real

on every CPU at every point inside or outside the announce window.

Symptom 2: that invariant relies on curr_tick and announce_remaining
moving in lock-step. The original loop advanced curr_tick at the top
of each outer iteration but only decremented announce_remaining after
the handler returned, with the lock dropped around the handler. A
non-announcing CPU calling sys_clock_tick_get() during a handler
therefore observed curr_tick + announce_remaining one batch (dt) into
the future, breaking the invariant for the duration of the handler.

Fix: pair the updates so they happen under the same lock acquisition
before the handler runs:

    curr_tick += t->dticks;
    announce_remaining -= t->dticks;

Symptom 3: commit d157b3da19 ("kernel: timeout: keep announce_remaining
stable across same-tick group") kept announce_remaining > 0 across the
inner drain of a same-tick group so that (a) the SMP early-return
would still detect "another CPU is announcing" and (b) elapsed() in
same-tick handlers would still anchor to the firing tick. With the
lock-state role moved to announcing_cpu, neither property depends on
announce_remaining anymore: any_cpu_announcing() carries (a) and
this_cpu_announcing() carries (b), both throughout the loop on the
announcing CPU regardless of the budget. The inner drain do-while
loop is therefore no longer needed -- the outer loop's
t->dticks <= announce_remaining naturally accepts dticks == 0
follow-on timeouts even when announce_remaining has reached 0.

Fixes #106317

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-20 10:55:56 +02:00
Nicolas Pitre
90d1963745 sys: util: move lowercase min/max/clamp to a new minmax.h
Since commit 37717b229f ("sys: util: rename Z_MIN Z_MAX Z_CLAMP to min
max and clamp"), <zephyr/sys/util.h> unconditionally defines function-
like macros named `min`, `max`, and `clamp` in the global namespace (in
C mode). util.h gets pulled in transitively by very broad headers,
including the POSIX layer's <pthread.h>, so any third-party C code that
uses these names as ordinary identifiers (e.g. XNNPACK's static `clamp`
helper and its public `clamp` struct field) fails to build as soon as
<pthread.h> is included.

Following the approach used by Linux, move the lowercase `min`, `max`,
`min3`, `max3`, and `clamp` macros (and their helpers) into a new
<zephyr/sys/minmax.h> header that has to be included explicitly by
source files that want them. util.h keeps the uppercase MIN/MAX/CLAMP,
so most code is unaffected; only the (much smaller) set of files that
actually use the lowercase variants needs to pick up the new include.

Fixes #107853.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-19 17:49:24 -04:00
Peter Mitsis
4424aa681e kernel: pipe: user threads may not re-init pipe
Updates z_vrfy_k_pipe_init() to use K_SYSCALL_OBJ_NEVER_INIT()
instead of K_SYSCALL_OBJ() to prevent a user thread from
re-initializing a pipe. This aligns the pipe initialization
behavior to that of other kernel objects such as message queues.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-05-15 23:28:30 +02:00
Li Jie
b0ed3dde1d spinlock: Validate support for up to 8 CPUs on 64-bit systems
Extends SPIN_VALIDATE feature to support up to 8 CPUs on
64-bit systems while maintaining backward compatibility with 32-bit
systems (which still support up to 4 CPUs).

Many modern SoCs have more than 4 CPU cores, yet the SPIN_VALIDATE
feature was only available for systems with MP_MAX_NUM_CPUS <= 4.
This limitation becomes increasingly relevant as multi-core designs
with 6-8 cores become common in both embedded and server applications.

The implementation leverages pointer alignment guarantees:
- On 32-bit systems: pointers are 4-byte aligned → 2 free bits →
  up to 4 CPUs
- On 64-bit systems: pointers are 8-byte aligned → 3 free bits →
  up to 8 CPUs

Signed-off-by: Li Jie <lijie.1996@picoheart.com>
2026-05-15 23:27:26 +02:00
Daniel Leung
4915839510 kernel: add kobj NULL check in k_thread_name_copy()
Inside k_thread_name_copy(), we call k_object_find() to find
the associated thread object of the incoming thread. However,
the finder can return NULL if incoming pointer address has
no kobj associated. So we need to check for NULL before
dereferencing k_object to look inside. Since k_object_find()
returns NULL if input object is NULL, there is no need to
specifically test thread pointer for NULL, and only need to
check for the return of k_object_find().

Signed-off-by: Daniel Leung <daniel.leung@intel.com>
2026-05-15 12:38:52 -05:00
Mayur Salve
f94efe2607 arch: riscv: address review nits on TLS canary patches
Add copyright year, reflow comments, remote duplicate __tls_end,
and remove commented-out code.

Signed-off-by: Mayur Salve <msalve@qti.qualcomm.com>
2026-05-14 21:52:56 +02:00
Mayur Salve
2d6bdab5dc arch/riscv: fix early TLS setup and save callee-saved registers
Fix the early TLS initialization sequence on RISC-V to correctly set up
the thread pointer (tp) before any C code runs, and save/restore
callee-saved registers as required by the ABI.

Signed-off-by: Mayur Salve <msalve@qti.qualcomm.com>
2026-05-14 21:52:56 +02:00
Mayur Salve
729110c12f arch: riscv: use TLS-based stack canary guard
This change enables per thread stack canary for RISC-V.

RISC-V GCC accesses the stack canary via a fixed offset from the
thread pointer (tp) when -mstack-protector-guard=tls is used. The
compiler emits code equivalent to:

  lw t0, 0(tp)   # load canary from tp+0

Additionally, tp is zeroed in arch_kernel_init() when TLS is enabled,
which means any C function called before thread setup completes (such
as z_early_rand_get or data_copy_xip_relocation) would fault trying
to access the canary.

Introduce STACK_CANARIES_TLS_PREPEND, which places the
.stack_chk.guard section at offset 0 of the TLS block, before .tdata
and .tbss. The compiler flags -mstack-protector-guard-reg=tp and
-mstack-protector-guard-offset=0 are passed so GCC generates the
correct canary access.

With STACK_CANARIES_TLS_PREPEND the per-thread TLS block layout is:

  tp --> +------------------+  offset 0
         | .stack_chk.guard |  (__stack_chk_guard)
         +------------------+
         | .tdata           |  (initialized TLS data)
         +------------------+
         | .tbss            |  (zero-initialized TLS data)
         +------------------+

The RISC-V reset path is extended to initialize tp before any C code
runs by allocating a TLS area on the boot stack and calling
arch_riscv_early_tls_stack_update(). Early boot functions that run
before tp is set up (z_early_rand_get, data_copy_xip_relocation) are
marked FUNC_NO_STACK_PROTECTOR to avoid canary access before tp is
valid.

Signed-off-by: Mayur Salve <msalve@qti.qualcomm.com>
2026-05-14 21:52:56 +02:00
Flavio Ceolin
fdc42fa256 kernel: userspace: fix SMP use-after-free
obj_list traversal held lists_lock, but removals held objfree_lock
(k_object_free) or obj_lock (unref_check). On SMP a concurrent
thread could free the node an iterator had saved as next.

Drop objfree_lock and require lists_lock for every obj_list
modification. k_object_free() now holds it across find+remove;
k_thread_perms_clear() takes it around unref_check().

Signed-off-by: Flavio Ceolin <flavio@hubblenetwork.com>
2026-05-13 09:14:14 +02:00
Jamie McCrae
1d935da700 kernel: Add support for dts RAM configuration
Allows using the chosen SRAM node for RAM configuration without
using Kconfig values

Signed-off-by: Jamie McCrae <jamie.mccrae@nordicsemi.no>
2026-05-11 08:45:38 +02:00
Peter Mitsis
8509270a0e kernel: poll timeout: Fix race condition
This fixes a subtle race condition in the poll timeout expiration
handler triggered_work_expiration_handler(). There was a small
window of opportunity between when sys_clock_announce() unlocks
interrupts and that the handler re-locked them that one or more
higher priority interrupts (or threads running on another CPU)
could abort the poll's timeout.

As each place where the timeout could be added or aborted already
locked the poll.c::lock spinlock, this commit updates the handler
to lock that spinlock upon entry and bail early if the timeout
has been detected to have been canceled.

It also changes the handler's call to k_work_submit_to_queue() to
z_work_submit_to_queue() since it is known that it will not be
rescheduling at that point within the timer ISR.

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-05-11 04:04:14 +02:00
Peter Mitsis
1b8c7a3038 kernel: thread timeout: Fix race condition
This fixes a subtle race condition in the thread timeout expiration
handler z_thread_timeout(). There was a small window of opportunity
between when sys_clock_announce() unlocked interrupts and that the
handler re-locked them that one or more higher priority interrupts
(or threads running on another CPU if in an SMP environment) could
abort the thread's timeout.

The fix has two parts. Part one ensures that _sched_spinlock is held
in every location before a thread's time can be canceled. Of the
various locations, only z_unpend_thread() was found to need updating.
Part two updates the timeout handler z_thread_timeout() to bail early
if the thread's timeout has been found to be canceled (or re-used)
during that aforementioned window.

Fixes #106653

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-05-11 04:04:02 +02:00
Christoph Busold
dad5096f85 drivers: entropy: Add support for architectural entropy drivers
Add new inline function entropy_get_default_device which returns
the "zephyr,entropy" device or the architectural entropy device,
if the former is not set, and use that in all places to query the
entropy device.

This allows using architectural drivers which do not have a DT
node.

Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
2026-05-06 07:05:12 +02:00
Nicolas Pitre
1296dc85e3 kernel: timer: make k_timer_start duration match documented semantics
The documentation for timer objects in
doc/kernel/services/timing/timers.rst has always stated:

    The timer's duration is a **minimum** delay relative to the time
    the timer was started.

However the implementation did not actually honour that contract. When
starting a relative duration, k_timer_start() was subtracting one from
duration.ticks before passing it to z_add_timeout(), cancelling out
z_add_timeout()'s conservative round-up. The net result was that
k_timer_start(K_TICKS(N), ...) could fire anywhere from just after 0
up to N ticks later -- strictly less than the documented minimum.

The in-tree comment acknowledged the mismatch ("i.e. k_timer_start()
doesn't treat its initial sleep argument the same way k_sleep() does,
but historical") and kept the subtraction for backwards compatibility.

This has lasted long enough. Drop the subtraction (and the companion
max(1, ...) whose only purpose was to keep the subtraction from
underflowing). k_timer_start() now honours its documented "minimum
delay" contract, matching the behaviour of k_sleep() for the same
tick count.

Callers that relied on the old "approximately N ticks" timing will
see up to one extra tick of delay on the initial fire, when the
call happens partway through a tick. Subsequent periodic fires are
unaffected: they are rescheduled from the timer ISR at an exact
tick boundary and continue to honour the period as before.

Note that a timer manually re-armed from within its own expiry
callback (rather than via the periodic 'period' argument) does not
suffer from the extra tick either: the callback runs inside
sys_clock_announce_locked(), so z_add_timeout()'s round-up is skipped
and the new fire lands at an exact tick stride. This preserves the
behaviour that the original -1 on the duration was presumably trying
to achieve in the first place, now obtained via the proper mechanism.

A few in-tree tests were tuned too tightly against the old
"approximately N" timing. Widen their tolerances to match the new
"at least N" contract:

  - tests/kernel/timer/timer_api: add one tick of slack in
    interval_check() to absorb the round-up.
  - tests/kernel/context: widen idle-timer slop by one tick.
  - tests/kernel/workq/work: express the busy-wait margin in ticks
    in the "running cancel" tests.
  - tests/kernel/threads/no-multithreading: on tickful kernels the
    pending IRQ delivered after irq_unlock()/k_cpu_idle() only
    announces one tick; wait one extra tick or loop idling until
    the timer callback runs.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-04 21:55:33 +02:00
Nicolas Pitre
2e2202af61 kernel: timeout: make z_add_timeout round-up conditional on announce
z_add_timeout() has always added one tick to the incoming tick count,
as a conservative round-up so that a request issued partway through a
tick still waits for "at least N full ticks" before the fire. This
round-up is correct in the general case, but *wrong* when the call
happens from within sys_clock_announce_locked() -- i.e. from a timer
expiration callback running at the tick-processing boundary where
elapsed() already returns 0. In that context there is no fractional
tick to compensate for, and the round-up simply makes every scheduled
timeout one tick late.

Two in-tree callers were already compensating for this caller-side:

* The k_timer periodic reschedule path in z_timer_expiration_handler
  always runs from inside sys_clock_announce_locked() and subtracted
  1 from the period before calling z_add_timeout(). Under
  CONFIG_TIMEOUT_64BIT, the same path additionally added +1 inside
  K_TIMEOUT_ABS_TICKS() to undo a related round-down.

* z_time_slice_reset() armed the slice timer with K_TICKS(slice_size
  - 1) so the resulting fire would land at exactly slice_size ticks.
  This one is reachable from both thread context (the +1 cancels the
  -1) and from announce context via update_cache() during a
  ready-thread wakeup (the +1 isn't applied, leaving the slice short
  by one tick). The latter is what actually trips
  tests/kernel/tickless/tickless_concept on every platform once the
  conditional below lands without dropping these workarounds.

All three are symptoms of the same root cause.

Handle it at the source: make the +1 conditional on announce_remaining
== 0. When scheduling from the timer ISR, we are already at a tick
boundary by construction, so no round-up is needed and periodic timers
now reschedule at exact period intervals without any caller-side
compensation. Drop the -1 in z_timer_expiration_handler's period
path, the +1 in its 64-bit absolute-reschedule companion, and the -1
in z_time_slice_reset(), since all three existed solely to paper over
this mismatch.

This change only affects timeouts scheduled from announce context
(periodic k_timers rescheduling themselves, callbacks starting new
timers, slice-timer rearm during ready-thread wakeup). All other
callers -- k_sleep(), z_abort_timeout(), initial k_timer_start() from
a thread, k_sched_time_slice_set() -- continue to use the +1 round-up.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-04 21:55:33 +02:00
Nicolas Pitre
d157b3da19 kernel: timeout: keep announce_remaining stable across same-tick group
When sys_clock_announce_locked() processes a tick that has multiple
timeouts queued for it, the timeout queue stores the second and
subsequent ones with dticks == 0 (relative to the first). The original
loop fired each in turn and decremented announce_remaining at the
bottom of every iteration:

    announce_remaining -= dt;

For the first timeout in a same-tick group, dt is the cumulative tick
delta that brought us to that tick. After that subtraction
announce_remaining can drop to zero, even though there are still
same-tick callbacks queued for the loop to fire. Each subsequent
same-tick callback then runs while announce_remaining == 0, which
breaks two invariants the rest of the kernel relies on:

* The SMP early-return at the top of sys_clock_announce_locked():

      if (IS_ENABLED(CONFIG_SMP) && (announce_remaining != 0)) {
          announce_remaining += ticks;
          k_spin_unlock(&timeout_lock, key);
          return;
      }

  is meant to detect that another CPU is already inside the loop and
  fold the new ticks into the ongoing announce. The lock is released
  around each callback, so during a same-tick callback another CPU
  can grab the lock, see announce_remaining == 0, miss the early
  return, set announce_remaining = ticks of its own, and start
  walking the queue in parallel with the original announcer -- the
  exact race the early return is supposed to prevent.

* The elapsed() helper:

      return announce_remaining == 0 ? sys_clock_elapsed() : 0U;

  is meant to return 0 for any z_add_timeout() / z_abort_timeout()
  call that happens from inside a tick-processing callback, so that
  timeouts scheduled from such a callback are anchored to the
  currently-firing tick rather than to a fresh sys_clock_elapsed()
  reading. With announce_remaining == 0 mid-loop, two callbacks on
  the same tick observe inconsistent semantics: the first one (that
  saw announce_remaining > 0) gets dticks anchored to the firing
  tick, while subsequent ones (seeing 0) get dticks computed against
  a fresh elapsed() reading and end up off by one tick. Two periodic
  timers that happen to fire on the same tick will therefore
  permanently drift apart by one tick going forward.

Restructure the loop so the announce_remaining decrement happens once
per tick rather than once per timeout: the outer while drives forward
across distinct ticks, and an inner do-while drains all timeouts
queued on the current tick before announce_remaining is updated.
announce_remaining now stays at its pre-tick value for the entire
same-tick group, which both the SMP early return and elapsed()
correctly observe as non-zero.

remove_timeout()'s dticks propagation is also unnecessary in this
loop because curr_tick is advanced by t->dticks before the timeout
is unlinked, which keeps the next item's stored dticks valid relative
to the new curr_tick. sys_dlist_remove() on its own is sufficient.

Loop structure suggested by Peter Mitsis.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-04 21:55:33 +02:00
Nicolas Pitre
603fa4818e drivers: timer: assert sys_clock lock held where required
Add sys_clock_is_locked(), the analog of z_spin_is_locked() for the
timer lock exposed via sys_clock_lock(). Use it to assert lock
ownership in sys_clock_set_timeout() and sys_clock_elapsed() of the
six timer drivers that were migrated to sys_clock_lock() and
consequently no longer acquire anything internally in those callbacks
(arm_arch_timer, riscv_machine_timer, xtensa_sys_timer, hpet,
apic_tsc, intel_adsp_timer).

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-05-01 11:18:04 -05:00