Commit graph zephyr/kernel
Author SHA1 Message Date
Anas Nashif
f1521edd22 kernel: move USE_SWITCH out of the SMP options menu
USE_SWITCH and USE_SWITCH_SUPPORTED describe a general context-switch
primitive (_arch_switch) that is usable on uniprocessor systems, not
just SMP configurations. Having them under the "SMP Options" menu in
kernel/smp/Kconfig was misleading.

Move both symbols to kernel/Kconfig proper, just before the SMP menu is
sourced. SMP still depends on USE_SWITCH; the guarding and arch selects
are unchanged.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-12 20:02:52 -04:00
Anas Nashif
cde0a961d8 kernel: thread: factor setup_thread_stack into helpers
setup_thread_stack() interleaved stack-geometry math, optional virtual
memory mapping, and several independent #ifdef'd steps (init fill,
sentinel, TLS/local-data/random headroom, stack info) into one body,
juggling five locals across a dozen preprocessor branches.

Introduce a small struct stack_geometry and split the work into focused
helpers: compute_stack_geometry(), map_thread_stack(), fill_init_stack(),
set_stack_sentinel(), reserve_stack_headroom() and set_stack_info().
Each optional step keeps its #ifdef internally and collapses to an
ARG_UNUSED no-op when disabled, so the change is cosmetic with no
runtime cost. setup_thread_stack() is now a short linear sequence and
the headroom accounting (which must stay ordered: TLS, then local data,
then randomization, then alignment) is isolated in one place.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-09 18:11:09 -04:00
Anas Nashif
1c47cb6eb4 kernel: thread: factor z_setup_new_thread config blocks into helpers
z_setup_new_thread() had grown into a ~140-line body dominated by
roughly fifteen #ifdef/#endif pairs interleaved with the actual setup
logic, making the initialization sequence hard to follow.

Push each optional-feature block down into its own static inline
helper (init_thread_obj_core, init_thread_userspace, init_thread_name,
add_thread_to_monitor, init_thread_usage, and so on). Each helper keeps
its #ifdef internally and collapses to an ARG_UNUSED no-op when its
feature is disabled, so the change is purely cosmetic with no runtime
cost. The function body now reads as a linear sequence of calls.

The only block that affects control flow, the
CONFIG_ARCH_HAS_CUSTOM_SWAP_TO_MAIN early return, is converted to an
IS_ENABLED() guard since _current and resource_pool always exist; it
dead-code-eliminates when the option is off. A no-op stub is added for
setup_shadow_stack so its call site needs no #ifdef either.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-09 18:11:09 -04:00
Nicolas Pitre
d33280510a kernel: timeout: add bucketed delta-list backend
Add a single-level bucketed delta list as a fourth selectable timeout
backend: a simpler relative of the timer wheel that keeps the wheel's
O(1) near-future insertion without its second tier, sift/defer
machinery, or tickless-idle penalty.

Timeouts expiring within the next CONFIG_TIMEOUT_BUCKET_LISTS ticks go
into a per-tick bucket list (O(1) insert and remove), tracked by an
occupancy bitmap so "ticks until next expiry" is O(1). Everything
beyond goes into one overflow list sorted by absolute expiry (O(n)
insert, O(1) remove, no delta fix-up). As curr_tick crosses the bucket
window, overflow entries that have come within range migrate into
buckets.

Like the other backends it is a single implementation header
(kernel/timeout_bucket.h) included only by timeout.c. Two properties
distinguish it from the other non-default backends, both confirmed by
running the full kernel timer/context/common suites with the backend
forced on (qemu_x86, x86_64, cortex_a53 SMP, riscv64):

  - It fits the generic next_gap/advance/pop_due announce loop -- no
    backend-owned announce. Migration is on demand against the actual
    next event, so an idle CPU with a distant timeout sleeps straight
    to it: it passes tests/kernel/context cpu_idle / timer_interrupts,
    which the wheel fails (the wheel wakes at least every 32 ticks).
  - Same-tick firing order stays FIFO (bucket and overflow inserts
    append; migration preserves order), so tests/kernel/common's
    timeout_order passes, unlike the min-heap.

The per-node representation is the delta list's node + dticks with no
extra field: a bucket entry stores its bucket index (< BUCKET_LISTS) in
dticks, an overflow entry its absolute expiry (>= BUCKET_LISTS); the
ranges never overlap, so it shares the dlist per-node helpers. Absolute
expiry needs 64-bit ticks, so the backend depends on TIMEOUT_64BIT. It
inherits the shared z_add_timeout round-up and the inflight_timeout
synchronization like the other backends. EXPERIMENTAL.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-09 18:02:12 -04:00
Nicolas Pitre
e0244d279b kernel: timeout: add timer-wheel backend
Add a hierarchical timer wheel as a third selectable timeout backend.
Timeouts are bucketed by expiry distance: one list per tick for the
next 32 ticks ("soon"), one list per 32-tick band for the next ~1024
("later"), and a sorted overflow list ("distant"). Insertion and
removal are O(1) for the common near-future case; every 32 ticks the
announce path sifts the current "later" band into "soon" and refills
it from "distant". This scales well when many short-lived timeouts are
pending.

Unlike the dlist and min-heap backends, the wheel does not fit the
generic next_gap/advance/pop_due announce primitives: its per-tick
advance is a bitmap scan that jumps over empty ticks, and its sift is
a time-driven event tied to no single timeout. Rather than contort
that (already subtle) state machine, the wheel uses the backend-owned
announce escape hatch: it defines _TIMEOUT_BACKEND_OWNS_ANNOUNCE and
implements z_timeout_q_announce(), which sys_clock_announce_locked()
calls in place of the generic loop.

Like the other backends the wheel is a single implementation header
(kernel/timeout_wheel.h) included only by timeout.c, so its state, its
operations and that announce loop reach the shared state (curr_tick,
announce_remaining, inflight_timeout) directly; nothing extra needs
exposing. The SMP re-entry guard, announcing_cpu and the reprogram
remain in timeout.c. The wheel fires handlers through the same
inflight_timeout dance, so the post-#109977 abort/in-flight
synchronization works unchanged, and it carries no per-node
ANNOUNCING/ABORTED sentinels.

struct _timeout grows a wheel-only flags field (which wheel tier a
timeout occupies), gated by CONFIG_TIMEOUT_BACKEND_WHEEL so the dlist
and heap builds are unaffected.

The backend is EXPERIMENTAL. Two known limitations, inherent to the
wheel algorithm (not the abstraction):
  - No same-tick firing-order guarantee (sifted timeouts are
    prepended), so the timeout_order test does not apply.
  - next_timeout() never exceeds 32 ticks because a sift is always
    pending, so the wheel wakes a tickless-idle CPU at least every 32
    ticks. This fails tests that assert zero spurious idle wakeups
    (tests/kernel/context cpu_idle / timer_interrupts) and is a power
    regression versus the dlist and heap backends.
Verified the timeout-functional suites (timer_api, timeout,
timepoints, sleep, sched/deadline) pass with the wheel on qemu_x86,
x86_64, cortex_a53 SMP and riscv64; dlist and heap remain unaffected.

The timer-wheel data structure, bucketing scheme and sift algorithm
are the work of Peter Mitsis (PR #108339), re-homed here behind the
timeout backend interface.

Co-authored-by: Peter Mitsis <peter.mitsis@intel.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-09 18:02:12 -04:00
Nicolas Pitre
5188d5009b kernel: timeout: add min-heap backend
Add a binary min-heap as a selectable timeout backend alongside the
sorted delta list, plugging into the backend abstraction rather than
forking timeout.c. Pending timeouts are kept in a min-heap keyed on
absolute expiry tick; insertion and arbitrary removal are O(log n),
which scales better than the delta list's O(n) insertion when many
timeouts are pending.

The backend is a new kernel/timeout_minheap.h: the heap instance (a
min_heap_ref from the previous commit), its comparator and the
z_timeout_q_*() operations, included only by timeout.c (after curr_tick)
so the whole backend stays private to that translation unit. struct
_timeout gains the abs_ticks + heap_handle representation, selected by
Kconfig (the delta list's node + dticks is the #else), with the common
fn pointer kept as a shared trailing member; the per-node helpers in
timeout_q.h grow a matching min-heap variant. The kernel shell thread
dump prints the backend's raw scheduling field (dticks or abs_ticks)
under #ifdef.

Because the shared z_add_timeout() already applies the post-#107452
conditional tick round-up, the heap inherits it: the backend's insert
simply stores abs_ticks = curr_tick + dticks. Likewise in-flight
handler synchronization (PR #109977) lives in timeout.c, so the heap
node carries no ANNOUNCING/ABORTED sentinels -- "not queued" is just
heap_handle.idx == 0.

The backend is EXPERIMENTAL, depends on TIMEOUT_64BIT (absolute ticks
need 64-bit precision), and uses a fixed-capacity heap
(CONFIG_TIMEOUT_HEAP_MAX_ENTRIES) whose overflow is a fatal error.

The min-heap algorithm, struct fields, Kconfig and capacity model are
derived from Sayooj K Karun's min-heap timeout subsystem (#106013),
reworked here to fit the pluggable backend and the current in-tree
in-flight-handler synchronization.

Co-authored-by: Sayooj K Karun <sayooj@aerlync.com>
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-09 18:02:12 -04:00
Nicolas Pitre
b381345043 kernel: timeout: factor delta-list queue into a pluggable backend
The timeout queue data structure and the front-end logic in timeout.c
are entangled through first()/next()/remove_timeout()/next_timeout()
and the open-coded insertion and announce loops. This makes it hard to
offer alternative queue implementations (min-heap, timer wheel) without
either forking timeout.c or threading #ifdefs through its trickiest
code, the announce path in particular.

Introduce a backend interface and move the sorted delta list behind it.
The split mirrors how the queue is actually used:

  - The per-node helpers z_init_timeout() / z_is_inactive_timeout()
    operate on a single struct _timeout and are needed tree-wide
    (timer.c, work.c, poll.c, ...), so they live in timeout_q.h and
    depend only on struct _timeout's fields.

  - The queue itself (its instance and the z_timeout_q_*() operations)
    is used only by timeout.c, so it becomes a backend implementation
    header (kernel/timeout_list.h) that timeout.c includes directly,
    after its shared state. The header is private to that translation
    unit; no other file sees the queue or its instance, so the instance
    is plain static and needs no extern.

timeout.c keeps all shared state (curr_tick, announce_remaining,
announcing_cpu, inflight_timeout, timeout_lock), the add/abort/query API
translation, the in-flight handler synchronization, and a single
announce loop shared by every future backend.

The backend exposes:
  - z_timeout_q_insert / _remove             add and arbitrary remove
  - z_timeout_q_remainder / _next_timeout    queries
  - z_timeout_q_next_gap / _advance / _pop_due   announce-loop primitives

The announce loop is recast in terms of the last three primitives:
find the gap to the next event, advance the backend (and curr_tick) to
it, then drain whatever is due. For the delta list this is behaviorally
identical to the previous open-coded loop: the residual advance reuses
the same head->dticks fixup, same-tick ties drain via successive
pop_due() calls, and announce_remaining is still decremented per tick
group while announcing_cpu carries the announcing state.

So the front end no longer reaches into the queue representation, three
checks are made backend-neutral: z_add_timeout()'s assert and
z_timer_expiration_handler()'s "restarted?" test now use
z_is_inactive_timeout() instead of sys_dnode_is_linked(), and
Z_TIMER_INITIALIZER zero-inits the queue fields (.fn only) rather than
naming .node/.dticks. struct _timeout is otherwise untouched; no
sentinel or per-node state changes.

The Kconfig backend choice and the divergent node layouts are deferred
until a second backend lands; with only the delta list present they
would be churn with no benefit.

Verified no behavioral change: tests/kernel/timer/{timer_api,timeout},
tests/kernel/common, and tests/kernel/sleep pass on qemu_x86,
qemu_x86_64, qemu_cortex_a53 (SMP) and qemu_riscv64 (exercising the
SMP announcing_cpu path and 64-bit dticks).

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-09 18:02:12 -04:00
Anas Nashif
1e265cf776 kernel: kheap: keep static heap init as a PRE_KERNEL_1 SYS_INIT
Static heap init was converted from SYS_INIT(PRE_KERNEL_1, OBJECTS) to a
K_KERNEL_INIT_PRE() hook. Those hooks run before any PRE_KERNEL_1
SYS_INIT, which broke heap KASAN: K_HEAP_KASAN_ENABLE() registers a
heap's shadow at (PRE_KERNEL_1, 0), relying on running before static
heap init. With the hook, sys_heap_init() ran first, heap->kasan_ba
stayed NULL and tracking was silently disabled. On the system heap this
let the tests/lib/heap_kasan overflow cases corrupt real chunk metadata
and panic with "heap corruption (free chunk linkage)".

Revert only the kheap portion of that conversion and keep static heap
init as SYS_INIT(PRE_KERNEL_1, CONFIG_KERNEL_INIT_PRIORITY_OBJECTS),
restoring the ordering the KASAN shadow registration depends on. The
mem_slab and mailbox conversions in the same series are unaffected and
stay on K_KERNEL_INIT_PRE. A comment records why kheap must remain a
SYS_INIT so it is not reconverted.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 19:42:44 -04:00
Anas Nashif
9a27e2e135 kernel: userspace: use kernel init hooks for mem domain and shmem
Convert the two userspace PRE_KERNEL_1 SYS_INIT users to
K_KERNEL_INIT_PRE() hooks:

  userspace.c   app_shmem_bss_zero  zero the BSS of app shared-memory
                                    regions
  mem_domain.c  init_mem_domain_module  set max partitions and build the
                                        default memory domain

This removes the last SYS_INIT users in kernel/, so kernel-internal init
no longer uses SYS_INIT at all.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Anas Nashif
ff6841547c kernel: obj_core: move kernel system object registration
The kernel object is a singleton (_kernel) whose object core type was
registered and linked from init.c. It is fully self-contained -- its
k_obj_type is referenced only by that registration -- so move it next to
the object core framework in obj_core.c, alongside the registration walk.

obj_core.c already includes kernel_internal.h (for the kernel init hook
and the kernel stats helpers) and reaches _kernel via kernel_structs.h, so
no new dependencies are introduced. CONFIG_OBJ_CORE_SYSTEM depends on
CONFIG_OBJ_CORE, so obj_core.c is always compiled when this block is built.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Anas Nashif
d78ced0f09 kernel: obj_core: register object-type init via kernel init hook
Convert the object core registration walk from SYS_INIT (PRE_KERNEL_1) to
a K_KERNEL_INIT_PRE() hook. This was the last SYS_INIT in kernel/*.c, so
kernel-internal initialization no longer uses SYS_INIT at all.

obj_core.c is linked whenever the object core framework is used (its
helpers are referenced by every object's init), so the walk runs whenever
CONFIG_OBJ_CORE is enabled, exactly as before -- only the registration
mechanism changes.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Anas Nashif
3a1d70f450 kernel: kheap, mem_slab, mailbox: use kernel init hooks
Convert the statically defined kernel objects' boot-time init from
SYS_INIT (PRE_KERNEL_1) to K_KERNEL_INIT_PRE() hooks:

  kheap     build the free block list for each static heap
  mem_slab  build the free block list for each static slab
  mailbox   build the pool of asynchronous message descriptors

These initializers live in their respective translation units, which are
linked only when the application references the corresponding APIs, so
each init runs only when its subsystem is actually used (pay-per-use),
matching the previous SYS_INIT behavior without the level/priority.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Anas Nashif
2d47a9b79c kernel: system_work_q: register init via kernel init hook
Replace the system work queue's SYS_INIT (POST_KERNEL) with a
K_KERNEL_INIT_POST() hook. The init function and its hook entry live in
this translation unit, which is linked only when something references the
system work queue (e.g. k_sys_work_q), so an image that never uses it
neither links this code nor runs this init -- the same pay-per-use
behavior SYS_INIT gave, now without the init level/priority.

The hook runs before POST_KERNEL device init so that drivers initialized
there may submit work to the queue.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Anas Nashif
6f268b6983 kernel: introduce kernel-internal init hooks
Add a lean, kernel-only initialization model to replace the kernel's
internal use of SYS_INIT. Kernel init order is fixed at compile time, so
these hooks drop the SYS_INIT level/priority machinery: a subsystem
registers a parameterless init function that the boot path runs at one of
two fixed phases.

  K_KERNEL_INIT_PRE(fn)   run in z_cstart() before PRE_KERNEL device init
  K_KERNEL_INIT_POST(fn)  run in bg_thread_main() before POST_KERNEL
                          device init

Each entry is a single function pointer placed, via an iterable section,
in the registering subsystem's own translation unit. The linker pulls
that unit (and the entry) into the image only when the subsystem is
otherwise referenced, so an init runs only when its subsystem is actually
linked. This preserves the pay-per-use linkage that SYS_INIT provided
while halving the per-entry size (one pointer versus init_entry's two) and
removing the level/priority sort. The boot-path walks reference only the
section bounds, so they never force a subsystem to be linked.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-06 15:10:55 -04:00
Guang Li
7a4c190291 kernel: extend cpu_mask and IPI support for up to 32 CPUs
Extend the cpu_mask field in struct _thread_base to support up to
32 CPUs by selecting uint32_t when CONFIG_MP_MAX_NUM_CPUS > 16.

Fix undefined behavior in IPI_ALL_CPUS_MASK where shifting 1 by
CONFIG_MP_MAX_NUM_CPUS (which can be 32) overflows uint32_t.
Replace with UINT32_MAX >> (32 - N) to produce the correct mask.

Signed-off-by: Guang Li <guang.li@jaguarmicro.com>
2026-07-06 15:10:03 -04:00
Anas Nashif
25249c3e23 kernel: timeout: guard arch_spin_relax() behind CONFIG_SMP
z_try_abort_timeout() can only set ret to -EAGAIN inside a
CONFIG_SMP-gated branch, so the trailing arch_spin_relax() call is
dead code on non-SMP builds. Optimized builds drop it, but a
coverage build (-O0) keeps the call and the linker then fails with
an undefined reference to arch_spin_relax: its weak definition
lives in idle.c, which is only compiled when CONFIG_MULTITHREADING
is set, so a no-multithreading + coverage build has no definition.

Wrap the check with IS_ENABLED(CONFIG_SMP) so the front end folds
the branch (and the arch_spin_relax reference) away even at -O0.
This is behavior-preserving since ret is never -EAGAIN without SMP.

The issue also triggers when building with llvm.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-05 15:32:21 -04:00
Anas Nashif
777ab58552 kernel: timeslicing: avoid recomputing the slice size on rearm
z_time_slice() computed z_time_slice_size(curr) to decide whether the
slice had expired, and then z_time_slice_reset() computed it a second
time to rearm the slice timeout. Cache the value instead: split the
rearm body into a slice_reset(slice_size) helper (z_time_slice_reset()
becomes a thin wrapper that still computes the size for its other
callers) and pass the already-known size from z_time_slice().

The size is still only computed when the slice has actually expired, so
the per-tick fast path is unchanged. In the CONFIG_TIMESLICE_PER_THREAD
case the expiry handler runs with the scheduler lock dropped and may
change the thread's slice configuration, so the cached value is
recomputed after the handler returns; behavior is therefore unchanged.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-02 12:59:47 -04:00
Anas Nashif
05ca560d2d kernel: thread: factor out thread name initialization
The strncpy + NUL-termination + arch_thread_name_set() hook sequence
behind CONFIG_THREAD_NAME was duplicated in z_impl_k_thread_name_set()
and z_setup_new_thread(). Extract it into a single set_thread_name()
helper and call it from both places, removing the duplicated logic and
the risk of the two copies drifting apart.

The helper treats a NULL name as "clear the name" (writes an empty
string), matching the previous z_setup_new_thread() behavior; the
k_thread_name_set() path always passed a non-NULL string, so its
behavior is unchanged.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-02 12:59:37 -04:00
Anas Nashif
c082e22f06 kernel: dynamic: combine thread state checks in stack free
z_impl_k_thread_stack_free() tested _THREAD_DUMMY and _THREAD_DEAD
with two separate z_is_thread_state_set() calls combined with a
logical OR. z_is_thread_state_set() returns (thread_state & state)
!= 0, so the two calls are equivalent to a single call with both
bits ORed into the mask. Collapse them into one call; this is
smaller and clearer with no change in behavior.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-02 12:59:28 -04:00
Anas Nashif
5e40ff3654 kernel: timeout: fix maybe-uninitialized warning in z_add_timeout
Initialize the variable at declaration to silence the false positive.
No functional change.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-02 12:58:00 -04:00
Anas Nashif
5967ceb85f kernel: init: register CPU object core via K_OBJ_TYPE_DEFINE
The CPU object core type was registered through a dedicated SYS_INIT. Like
threads, CPUs have no statically defined instances to walk at boot: each
CPU links its own object core in z_init_cpu(). The boot init only
registered the type and its stats descriptor.

Use K_OBJ_TYPE_DEFINE_TYPE_ONLY() for the CPU type, dropping
init_cpu_obj_core_list() and its SYS_INIT. The type is still registered at
PRE_KERNEL_1 by the single object core init walk, before z_init_cpu(0)
runs during prepare_multithreading(), so ordering is unchanged. The kernel
system object keeps its own init, as it links the singleton _kernel object.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-01 23:45:35 -04:00
Anas Nashif
ce9584adf7 kernel: thread: integrate object core via K_OBJ_TYPE_DEFINE
Threads registered their object core type through a dedicated SYS_INIT.
Unlike the other object types they have no statically defined instances
to walk at boot: each thread links its own object core as it is created
in z_setup_new_thread(). The boot init therefore only registered the type
and its stats descriptor.

Add a K_OBJ_TYPE_DEFINE_TYPE_ONLY() variant that registers the type (and
optional stats descriptor) without walking a static object section, and
use it for threads. The type is still registered at PRE_KERNEL_1 via the
single object core init walk, before any thread is created, so ordering is
unchanged. Only the internal cpu/kernel system objects (which are not in
iterable sections) now remain outside the table.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-01 23:45:35 -04:00
Anas Nashif
cfb9ad7d7b kernel: mem_slab: integrate object core via K_OBJ_TYPE_DEFINE
mem_slab was the one statically defined object type left out of the
object core registration table, because its boot-time init bundled two
unrelated concerns: building each slab's free block list (create_free_list,
mandatory functional init) and its object core duties (type init, stats
descriptor init, linking and per-slab stats registration).

Separate the two. The object core duties now go through the registration
table like every other object type, and the per-slab statistics buffer is
registered generically by the table walk: the descriptor optionally carries
the offset and size of an embedded stats buffer (only present under
CONFIG_OBJ_CORE_STATS), set via the new K_OBJ_TYPE_DEFINE_STATS() macro.
K_OBJ_TYPE_DEFINE() forwards to it with no per-object stats.

create_free_list stays as its own SYS_INIT: it is required whether or not
the object core framework is enabled (CONFIG_OBJ_CORE is off by default),
so it cannot move into the object-core-gated table walk. mem_slab.c is
linked only when slab APIs are used, so this init carries no extra
footprint for builds that do not use slabs.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-01 23:45:35 -04:00
Anas Nashif
21a9723208 kernel: obj_core: register object types from a single init point
Each kernel object type that participates in the object core framework
previously supplied its own SYS_INIT routine to initialize its
k_obj_type and to walk its static-object linker section, linking each
object core. These routines were near-identical across 11 object types
and differed only in the type id, the object struct and the obj_core
offset.

Replace that per-type boilerplate with a declarative K_OBJ_TYPE_DEFINE()
macro that emits a const descriptor into a new iterable ROM section, and
walk those descriptors once from a single SYS_INIT in obj_core.c. The
descriptor captures the type storage, type id, obj_core offset, the
static object section bounds and the object stride, which is all the
shared loop needs to initialize and link every statically defined
object.

Converted: condvar, event, fifo, lifo, mailbox, msgq, mutex, pipe, sem,
stack and timer. The non-uniform initializers (thread, mem_slab and the
internal cpu/kernel objects) are left unchanged and will be dealt with
in followup commits.

Footprint on qemu_cortex_m3 (tests/kernel/obj_core/obj_core, all types
enabled): flash 32012 -> 30880 (-1132 B), RAM unchanged. With
CONFIG_OBJ_CORE disabled the image is byte-for-byte identical.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-01 23:45:35 -04:00
Nicolas Pitre
c33571b372 kernel: timeout: cap the timeout passed to drivers centrally
sys_clock_announce() reports ticks elapsed since the previous announce. The
quantity that must stay in range is therefore the distance of the next
timeout from that announce, not its distance from "now" as next_timeout()
previously bounded. When an early timeout is repeatedly replaced by a later
one, announces are deferred and that distance can exceed the range even
though each set_timeout() delay was in range, which every driver has had to
defend against on its own.

Cap it in the core instead: SYS_CLOCK_MAX_WAIT becomes UINT32_MAX/2 (the
clamp, with the upper half of the range left as slack for a late announce)
and next_timeout() clamps the inter-announce distance to it. Drivers no
longer need to clamp for the announce range and only have to honour their
own cycle-count limits.

The core no longer emits K_TICKS_FOREVER; the reported delay is always
finite, saturating at SYS_CLOCK_MAX_WAIT. Drivers that stop the timer
entirely when idle now key off ticks == SYS_CLOCK_MAX_WAIT under
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE.

z_get_next_timeout_expiry() still reports to the idle and PM paths as a
signed int32_t. It keeps its K_TICKS_FOREVER default and adopts the
next_timeout() value only when that fits, so the unsigned cap can never
surface there as a negative number.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-01 23:44:31 -04:00
Nicolas Pitre
f4a2a27071 kernel: timeout: make the system clock tick interface unsigned
The sys_clock_set_timeout(), sys_clock_announce() and
sys_clock_announce_locked() interfaces carry a number of ticks to be
scheduled or announced. Those ticks have no negative meaning, so a signed
argument is both wasteful and error prone: it invites incorrect handling in
drivers and it halves the range that sys_clock_announce() can represent.

Switch the tick argument of these interfaces, along with the internal
announce_remaining accounting, from int32_t to uint32_t. This is a
mechanical change: driver bodies perform plain arithmetic on the value and
are unaffected by the signedness, and the kernel never passes a negative
tick count. The freed sign bit doubles the representable announce range,
which a later change will use to move the announce range limit out of the
drivers and into the core.

One config is not strictly neutral: under CONFIG_TIMEOUT_64BIT,
K_TICKS_FOREVER is a 64-bit value that an unsigned 32-bit argument can no
longer compare equal to, so the K_TICKS_FOREVER tests in drivers stop
matching there. K_TICKS_FOREVER is a k_ticks_t concept and has no business
in this tick count interface; a follow-up removes it from the driver side
entirely, which closes this gap.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-07-01 23:44:31 -04:00
Anas Nashif
9722f39203 kernel: fix doxygen warnings exposed by adding kernel/ to docs INPUT
Bringing the kernel/ sources into the Doxygen requirements-traceability
INPUT (for @satisfies links) surfaces a handful of latent documentation
warnings that would break the -W docs build. Fix them.

With these, the kernel/ tree is clean under doxygen 1.17.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-07-01 17:05:28 -04:00
Jinming Zhao
217e2913a1 lib: heap: kasan: lightweight heap write sanitizer for sys_heap
Lightweight write sanitizer for sys_heap.  Uses
-fsanitize=kernel-address compiler instrumentation combined with a
per-heap shadow bitarray (one bit per granule) to detect buffer
overflows, underflows, and use-after-free on write accesses.

Ships its own lightweight sanitizer runtime (__asan_store* callbacks);
does not depend on an external ASAN library and supports debugging on
real embedded targets.

Instrumentation is opt-in per CMake target via
zephyr_target_enable_heap_kasan(), or per directory via
zephyr_heap_kasan_enable_directory().  Only writes are checked
(-asan-instrument-reads=0).  Bulk-write library calls (memset,
memcpy, str*, printf family) are redirected to checked wrappers at
compile time via -Dfoo=__asan_foo, requiring no source changes in
application code.

Heap tracking is likewise opt-in: register each heap with
SYS_HEAP_KASAN_ENABLE() / K_HEAP_KASAN_ENABLE(), or enable
CONFIG_SYS_HEAP_KASAN_MALLOC / CONFIG_SYS_HEAP_KASAN_SYSTEM for the
common libc malloc and kernel system heaps.

Usage:

  CONFIG_SYS_HEAP_KASAN=y
  CONFIG_SYS_HEAP_KASAN_MALLOC=y   # auto-track malloc/free
  CONFIG_SYS_HEAP_KASAN_SYSTEM=y   # auto-track k_malloc/k_free

  # Instrument all sources of <target> (CMakeLists.txt)
  zephyr_target_enable_heap_kasan(app)
  # Or instrument sources under <dir>
  zephyr_heap_kasan_enable_directory(src/mymodule)

  /* Opt-in tracking for a custom heap */
  K_HEAP_DEFINE(my_heap, 4096);
  K_HEAP_KASAN_ENABLE(my_heap, 4096);

Signed-off-by: Jinming Zhao <jinmzhao@qti.qualcomm.com>
2026-07-01 05:21:22 -04:00
Tom Burdick
3096d819a9 kernel: De-duplicate symbol names
Some private symbol names were widely duplicated throughout the kernel
such as lock and handle_poll_event. This arguably made the kernel less
readable as the locality of the name lock is highly confusing.

Rename all compilation unit locks to match their usage (e.g.
mutex_lock). Improving readability.

Furthermore, by deduplicating these symbols we enable potential
amalgamation builds of the kernel where all C files are merged
into one large C file or compliation unit allowing for better
compiler visibility and optimization.

Signed-off-by: Tom Burdick <thomas.burdick@infineon.com>
2026-06-30 11:53:41 -04:00
Zach Van Camp
7e9d3e3476 kernel: Add RESET_ON_FATAL_ERROR kconfig option
Add an optional reset when a hard fault occurs.

Signed-off-by: Zach Van Camp <zach.vancamp@etcconnect.com>
2026-06-29 22:18:47 -04:00
Michael Zimmermann
480d0eee45 kernel: work: deprecate thread field in struct k_work_q
The `thread_id` member should be used instead, because this also works for
work queues which run on a thread, that was not started by the work queue
implementation.

Signed-off-by: Michael Zimmermann <michael.zimmermann@sevenlab.de>
2026-06-29 22:17:45 -04:00
Michael Zimmermann
e10af22991 tree-wide: replace usages of k_work_q thread with thread_id
The `thread` field is not initialized, if it's being animated via
k_work_queue_run instead of k_work_queue_start.

Signed-off-by: Michael Zimmermann <michael.zimmermann@sevenlab.de>
2026-06-29 22:17:45 -04:00
Nicolas Pitre
d7059e67e8 kernel: sched: ready unpended threads unconditionally in *unpend_all
z_unpend_all() and unpend_all() skipped readying a woken thread when
z_try_abort_thread_timeout() returned -EAGAIN, on the assumption that
an in-flight timeout handler on another CPU would ready the thread
itself once the scheduler lock was dropped.

That assumption no longer holds. z_thread_timeout() was changed to bail
out when it observes the timeout has been marked superseded, and
z_try_abort_thread_timeout() sets exactly that superseded mark on the
-EAGAIN path. So on SMP, when a waiter's timeout fires on one CPU at
the same moment *unpend_all() processes it on another:

  - *unpend_all() unpends the thread and, seeing -EAGAIN, does NOT
    ready it.
  - the in-flight z_thread_timeout() finds the timeout superseded and
    bails, so it does NOT ready it either.

The thread ends up off every wait queue and on no run queue: orphaned,
with no remaining wake source. (k_heap / sys_mempool waiters with a
finite timeout are the reachable callers.)

Fix: ready the thread unconditionally after aborting its timeout, the
same idiom z_unpend_first_thread_locked() and kernel/events.c already
use. The abort still marks the in-flight handler superseded so it bails
and cannot double-wake, and ready_thread() is idempotent regardless.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-26 22:29:48 -04:00
Appana Durga Kedareswara rao
940aaaed2a kernel: fatal: add arch hook pair for IRQ management during coredump
z_fatal_error() locks interrupts before calling coredump(). Some
backends (e.g. a UDP socket backend) need IRQ delivery to complete
their TX path. Introduce two weak arch hooks that bracket the
coredump() call:

  arch_coredump_fatal_irq_unlock(key, cookie) -- re-enables IRQs
  arch_coredump_fatal_irq_lock(cookie)        -- restores the mask

Both are declared in kernel_arch_interface.h under
CONFIG_DEBUG_COREDUMP_FATAL_UNLOCK_IRQS. The default weak
implementations in subsys/debug/coredump/ delegate to the standard
arch_irq_unlock()/arch_irq_lock() pair. Architectures where exception
entry independently masks IRQs (e.g. ARM64 DAIF.I) must override them.

Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
2026-06-25 22:18:57 +01:00
Peter Mitsis
23040b6b55 kernel: Inline z_sched_wake()
Inlining z_sched_wake() gives the thread_metric synchronization
benchmark an almost 40% boost. This restores it to its previous
levels when k_sem_give() was using z_unpend_first_thread().

Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
2026-06-25 06:19:19 -04:00
Mathieu Choplain
1a72685d23 kernel: events: update comment explaining the wake algorithm
The comment made it seem like we do two passes on the "threads that need
wakeup" list: a first one to unpend and another to ready. In reality,
we do only one path and perform both operations at one using
`z_sched_wake_thread_locked()`.

Update the comment explaining the algorithm so it reflects more
accurately what the code actually does.

Signed-off-by: Mathieu Choplain <mathieu.choplain-ext@st.com>
2026-06-24 11:45:57 +02:00
Anas Nashif
69824a8eb4 kernel: fix abbreviated stack safety check with stack sentinel
k_thread_runtime_stack_safety_threshold_check() passed the thread's
unused stack threshold directly as the scan window size to
z_stack_space_get(). When CONFIG_STACK_SENTINEL is enabled,
z_stack_space_get() reserves the first 4 bytes of the stack buffer for
the sentinel and shrinks its scan window by that amount. As a result a
fully unused window reported "threshold - 4" unused bytes, which is
always less than the threshold, so the handler fired on every call
regardless of the actual unused stack space, making the abbreviated
check useless on sentinel-based platforms.

Grow the requested scan window by the sentinel reservation so that up to
"threshold" bytes of the usable stack are actually examined. The
firing condition is unchanged on platforms without the sentinel.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-23 21:17:03 -04:00
Anas Nashif
0f7b0ba0f1 kernel: fix thread runtime stack safety syscall names
The runtime stack safety threshold syscalls were declared in kernel.h as
k_thread_runtime_stack_unused_threshold_{pct_set,set,get}, but their
implementations and verifiers in thread.c were named with an extra
"safety_" infix (k_thread_runtime_stack_safety_unused_threshold_*). Since
the __syscall declarations drive code generation, the generated _mrsh.c
files and the expected z_impl_/z_vrfy_ symbols never matched the
definitions, so the feature failed to build/link with
CONFIG_THREAD_RUNTIME_STACK_SAFETY and CONFIG_USERSPACE enabled.

Rename the implementations and verifiers to match the public API names
and fix the corresponding generated-header includes.

Also remove a duplicate, malformed z_vrfy for the threshold "get"
syscall that returned int instead of size_t and called the "set"
implementation with the wrong number of arguments. The correct verifier
is already defined alongside the other threshold syscalls.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-23 21:17:03 -04:00
Anas Nashif
dc61457251 kernel: stack canaries: disable TLS canaries for Clang on RISC-V
Clang's RISC-V backend rejects the -mstack-protector-guard=tls family of
flags used to anchor the stack canary to the TLS base register. With those
flags stripped, Clang falls back to an absolute reference to the
thread-local __stack_chk_guard symbol, which resolves to address 0. The
build compiles cleanly but the resulting image faults in the prologue of
the first canary-instrumented function (vprintk during early boot, before
the console is up), then re-faults forever in the trap handler, producing
no output at all.

Guard STACK_CANARIES_TLS so it cannot be selected for Clang on RISC-V,
falling back to the global canary which Clang handles correctly, and
exclude the kernel.memory_protection.stackprot_tls test from zephyr/llvm.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-23 17:40:13 -04:00
Nicolas Pitre
97fac4b782 kernel: mmu: remove LINKER_USE_PINNED_SECTION and __pinned_* tagging
CONFIG_LINKER_USE_PINNED_SECTION is the second half of the selective
kernel-pinning model removed in issue #108773. With the kernel image
now always resident at boot (previous commit), the __pinned_*
attribute family is a no-op: every page they would have segregated is
already pinned by z_mem_manage_init()'s whole-image loop, so the
tagging contract neither adds safety nor remains maintainable.

Drop it.

Mechanical removals:

* All ~219 in-tree uses of __pinned_text, __pinned_rodata,
  __pinned_data, __pinned_bss, __pinned_noinit, and __pinned_func
  across arch/x86, drivers/interrupt_controller, drivers/timer,
  arch/common, kernel, lib/libc, subsys/portability/posix, tests, and
  the syscall code generator (scripts/build/gen_syscalls.py).

* The assembly aliases PINNED_TEXT/RODATA/DATA/BSS/NOINIT used in
  arch/x86/core/ia32/*.S and drivers/interrupt_controller/
  intc_loapic_spurious.S become plain TEXT/RODATA/DATA/BSS/NOINIT.

* K_KERNEL_PINNED_STACK_DEFINE, K_KERNEL_PINNED_STACK_ARRAY_DEFINE,
  K_KERNEL_PINNED_STACK_ARRAY_DECLARE, K_THREAD_PINNED_STACK_DEFINE,
  and K_THREAD_PINNED_STACK_ARRAY_DEFINE are removed. The few
  in-tree callers (kernel/init.c, arch/arm/core/cortex_a_r/smp.c,
  arch/arm64/core/fatal.c, arch/rx/core/prep_c.c,
  arch/x86/core/prep_c.c, kernel/include/kernel_internal.h,
  tests/bluetooth/hci_uart_async) move to the corresponding
  non-pinned macros.

Machinery removals:

* Kconfig.zephyr drops CONFIG_LINKER_USE_PINNED_SECTION.
  qemu_x86_tiny and qemu_x86_atom_virt drop their =y overrides.

* include/zephyr/linker/section_tags.h drops the __pinned_* macro
  definitions (both arms). __isr collapses to an empty macro since
  its only purpose was to alias __pinned_func.

* include/zephyr/linker/sections.h drops PINNED_TEXT_SECTION_NAME,
  PINNED_BSS_SECTION_NAME, etc. and the bare PINNED_TEXT/RODATA/etc.
  forwarders, plus the _APP_SMEM_PINNED_SECTION_NAME constant.

* include/zephyr/linker/linker-defs.h drops the lnkr_pinned_*
  externs, the _app_smem_pinned_* externs, and the lnkr_is_pinned()
  / lnkr_is_region_pinned() inline helpers.

* include/zephyr/linker/utils.h drops the lnkr_pinned_rodata branch
  in linker_is_in_rodata().

* include/zephyr/linker/app_smem_pinned{,_aligned,_unaligned}.ld
  are deleted; cmake/linker/ld/target_configure.cmake stops
  configuring them.

* boards/qemu/x86/qemu_x86_tiny.ld and
  include/zephyr/arch/x86/ia32/linker.ld drop their pinned-section
  blocks and the now-redundant #ifndef CONFIG_LINKER_USE_PINNED_SECTION
  conditionals throughout the body. The
  LIB_KERNEL_IN_SECT / LIB_ARCH_X86_IN_SECT / LIB_ZEPHYR_IN_SECT /
  LIB_C_IN_SECT / LIB_DRIVERS_IN_SECT / LIB_SUBSYS_LOGGING_IN_SECT /
  LIB_ZEPHYR_OBJECT_FILE_IN_SECT / ZEPHYR_KERNEL_FUNCS_IN_SECT macros
  in qemu_x86_tiny.ld are deleted; they existed only to feed the
  pinned text/rodata/data/bss/noinit sections.

* kernel/mmu.c drops the mark_linker_section_pinned(lnkr_pinned_start,
  ...) call. The mark_linker_section_pinned() helper survives but is
  now gated only on CONFIG_LINKER_USE_BOOT_SECTION.

* arch/common/init.c and include/zephyr/arch/common/init.h drop
  arch_bss_zero_pinned(); arch/x86/core/ia32/crt0.S drops the call
  to it.

* arch/x86/core/userspace.c drops the eager k_mem_page_in() of the
  thread's privileged stack on user-mode entry. With the kernel
  image fully resident the stack is already mapped.

* arch/x86/gen_mmu.py drops map_region("lnkr_pinned") and the
  set_region_perms() calls for lnkr_pinned_text / lnkr_pinned_rodata.

* CMakeLists.txt drops the LINKER_USE_PINNED_SECTION block that
  generated APP_SMEM_PINNED_* variables and the
  pinned_partitions target property feeding gen_app_partitions.py.
  cmake/modules/extensions.cmake removes the PINNED_RODATA /
  PINNED_RAM_SECTIONS / PINNED_DATA_SECTIONS zephyr_linker_sources()
  location keywords and their snippet files.
  scripts/build/gen_app_partitions.py drops --pinoutput /
  --pinpartitions arguments and the pinned-output branch.
  subsys/testsuite/coverage/CMakeLists.txt drops its
  CONFIG_DEMAND_PAGING-conditional fork.

* scripts/build/gen_kobject_list.py drops the
  app_smem_pinned_start / _end fallback for kobject placement
  validation.

* tests/arch/x86/pagetables and tests/kernel/mem_protect/userspace
  drop their lnkr_pinned_text / lnkr_pinned_rodata branches.

* include/zephyr/arch/x86/ia32/arch.h folds IRQSTUBS_TEXT_SECTION
  to the unconditional ".text.irqstubs" form.

* tests/subsys/llext/src/syscalls_ext.c drops a stale comment about
  syscalls landing in .pinned_text.

Targeted retentions:

* arch/x86/core/bootargs.c keeps multiboot_cmdline and efi_bootargs
  in .noinit (was __pinned_noinit, which decayed to __noinit when
  LINKER_USE_PINNED_SECTION was unset). The multiboot and zefi loader
  paths write these buffers before Zephyr's BSS-zero step, so
  zeroing them at boot loses the cmdline.

* arch/x86/core/ia32/fatal.c keeps _df_esf and _df_stack in .noinit.
  They are scratch space written by the double-fault handler and have
  no zero-init requirement; keeping them in .noinit also preserves
  the historical post-noinit alignment that gen_mmu.py relies on
  (z_mapped_size is computed before CMake-injected iterable sections
  are appended to the linker script, so the post-noinit page padding
  is what keeps those sections within the mapped region).

* include/zephyr/arch/x86/ia32/syscall.h and
  include/zephyr/arch/x86/arch.h wrap the per-arch
  arch_syscall_invoke* / arch_is_user_context / arch_k_cycle_get_*
  implementations in @cond INTERNAL_HIDDEN. The public Doxygen
  contract lives on the prototypes in
  include/zephyr/arch/arch_interface.h; the per-arch implementations
  are internal. Without this, removing the __pinned_func attribute
  exposes the implementations to the doxygen-coverage delta check
  as 10 newly-undocumented APIs.

Documentation updates are deferred to a separate commit.

Issue: #108773

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-23 09:11:08 -04:00
Nicolas Pitre
86c6923e12 kernel: mmu: remove LINKER_GENERIC_SECTIONS_PRESENT_AT_BOOT
With qemu_x86_tiny moved off the =n side of this knob in the previous
commit, no in-tree configuration sets
CONFIG_LINKER_GENERIC_SECTIONS_PRESENT_AT_BOOT=n anymore. The mode it
selected is also unsafe by construction (issue #108773), so there is
nothing to deprecate -- drop the Kconfig symbol and every #ifdef on
it.

Consequences:

* z_mem_manage_init() pins the whole Zephyr image unconditionally and
  no longer calls arch_bss_zero() after demand-paging init. The
  pinning loop's TODO about "we will need linker regions for a subset
  of kernel code/data pages which are pinned in memory and may not be
  evicted" is replaced with a brief note that the kernel image is
  always resident now and that pageable kernel regions go through
  __ondemand_*.

* arch/x86/core/ia32/crt0.S always calls arch_bss_zero() once
  arch_bss_zero_boot and arch_bss_zero_pinned have run.

* kernel/kheap.c and kernel/userspace/userspace.c drop the
  pre-kernel/post-kernel "skip non-pinned" dance and zero/init each
  heap and app-shmem partition unconditionally at PRE_KERNEL_1.

* arch/x86/core/userspace.c drops the eager k_mem_page_in() of the
  thread's privileged stack before dropping to user mode -- with the
  full kernel image resident the stack is already mapped.

* arch/x86/gen_mmu.py always maps the Zephyr image with FLAG_P set
  (and lnkr_boot_* / lnkr_pinned_* regions just inherit that).

* CMakeLists.txt no longer force-pins z_libc_partition into
  pinned_partitions for app_smem.

* The EVICTION_LRU gate
  "depends on LINKER_GENERIC_SECTIONS_PRESENT_AT_BOOT" goes away;
  LRU is now safe by construction on any DEMAND_PAGING board and
  becomes the default whenever ARCH_SUPPORTS_EVICTION_TRACKING.

* DEMAND_PAGING_PAGE_FRAMES_RESERVE loses its conditional default of
  32 frames; the new model needs no reserve.

* boards/qemu/x86/qemu_x86_tiny.ld drops the FLASH MEMORY region and
  the flash_load_offset block that arranged the demand-paged
  generic-section layout. boards/qemu/x86/board.cmake drops the
  8 MB-RAM bump and the --map flash gen_mmu argument that paired with
  that mode.

* tests/arch/x86/pagetables and tests/kernel/fatal/exception drop the
  code paths that were guarded on =n.

* tests/kernel/mem_protect/demand_paging/mem_map.lru replaces its
  CONFIG_LINKER_GENERIC_SECTIONS_PRESENT_AT_BOOT=y override with an
  explicit CONFIG_EVICTION_LRU=y, since the symbol it relied on is
  gone.

The __pinned_* tagging convention still expands to its actual linker
sections under CONFIG_LINKER_USE_PINNED_SECTION; that symbol and the
~219 in-tree __pinned_* annotations are cleaned up in the following
commit.

Issue: #108773

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-23 09:11:08 -04:00
Christoph Busold
9828037e91 arch: riscv: Detach DEP config from EXECUTE_XOR_WRITE
Move EXECUTE_XOR_WRITE back into userspace and update the comment
to reflect that it only applies to memory partitions.

PMP_DATA_EXECUTION_PREVENTION is now independent from it and no
longer a requirement.

For full protection both options should be enabled. Otherwise,
some data areas my be left executable and writable at the same
time.

Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
2026-06-19 22:37:17 +02:00
Christoph Busold
70df4ac429 arch: riscv: Add CONFIG_EXECUTE_XOR_WRITE support using PMP
CONFIG_EXECUTE_XOR_WRITE is a basic countermeasure to prevent code
injection attacks by making data non-executable. This can be done
on RISC-V with PMP by removing the executable permission.

Since PMP slots are statically prioritized, we can add a region
covering the whole address space into the last register giving
only read-write access. This will be overridden by any other
configured register with lower index so that the code region
remains executable.

In order to make this register apply to kernel threads in M mode,
it needs to be locked, but this means we cannot remove it anymore.
For CONFIG_USERSPACE we therefore need to reserve another register
so that we can add another address space-wide region with higher
priority without any permission to override the last one and limit
user threads to their assigned regions. This needs to be removed
when switching from user to kernel threads, which is done in the
new function z_riscv_pmp_usermode_disable.

This feature simplifies CONFIG_PMP_KERNEL_MODE_DYNAMIC because we
already have a catch-all register for kernel threads.

Does not work with CONFIG_CODE_DATA_RELOCATION for now and requires
PMP region locking.

Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
2026-06-19 22:37:17 +02:00
Anas Nashif
bd1828652d kernel: thread: fix thread_obj_validate oops
thread_obj_validate()'s default case passed the (non-zero) error code
`ret` straight to K_SYSCALL_VERIFY_MSG(), which verifies its argument is
true. A non-zero `ret` therefore reads as "verified OK", so K_OOPS()
never raised and control fell through to CODE_UNREACHABLE. With GCC this
path isn't normally reached and __builtin_unreachable() is a no-op; with
Clang it traps (unimp), turning an access-denied into an illegal
instruction. Verify `ret == 0` so the oops is actually raised.

Signed-off-by: Anas Nashif <anas.nashif@intel.com>
2026-06-17 10:09:44 -04:00
Nicolas Pitre
1e68351a25 kernel: userspace: abort timer before freeing
Re-implementation of the fix originally proposed as PR #109426 on
top of the z_try_abort_timeout() infrastructure.

When a dynamically allocated k_timer's reference count drops to
zero its storage is about to be freed. If the timer's expiration
handler is in flight on another CPU at this moment, it will
dereference the freed memory.

Introduce k_timer_cleanup(): cancels the timeout and waits for any
in-flight handler to complete. Unlike k_timer_stop() it does not
run the user stop_fn nor wake pending threads -- the caller is
about to drop the storage and there is no further consumer. The
retry-on-EAGAIN loop around z_try_abort_timeout() guarantees that
when this returns 0 the handler has fully run; no in-band sentinel
is involved.

unref_check() in kernel/userspace/userspace.c calls
k_timer_cleanup() for K_OBJ_TIMER, guarded by K_OBJ_FLAG_INITIALIZED
(z_waitq_head() on an uninitialized dnode would otherwise read
garbage).

Add sys_port_trace_k_timer_cleanup_{enter,exit} hooks across the
tracing backends (default no-op).

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
016f135c9b kernel: timeout: remove the TIMEOUT_DTICKS_ABORTED sentinel
The only reader was z_is_aborted_thread_timeout() in kernel/sleep.c,
used to distinguish a normal timeout-driven wakeup from one caused by
k_wakeup(). That distinction can be made entirely from the time
remaining: a normal wakeup leaves left_ticks <= 0, while an early
k_wakeup() leaves a positive remainder. The signed_left > 0 check
that already follows handles both cases correctly. The
z_is_aborted_thread_timeout() check was an early-return optimization
that the time computation makes redundant.

Drop the early-return in z_tick_sleep(), remove z_is_aborted_timeout(),
z_is_aborted_thread_timeout(), and the TIMEOUT_DTICKS_ABORTED macro.
Stop writing dticks = ABORTED in z_try_abort_timeout()'s linked-removal
path; the same-CPU IRQ branch is simplified to a flat -EINVAL fall-
through (the comment is moved to the cross-CPU/-EAGAIN side which is
the only case the caller needs to distinguish).

After this commit, dticks is purely a queue-management delta; there
are no more in-band sentinel values.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
414c3cdeaf kernel: timeout: remove the dticks-based cancel mechanism
With every caller now using z_try_abort_timeout() and every handler
relying on subsystem-local wake-ownership flags (queue->finished,
K_WORK_DELAYED_BIT, poller.is_polling, killed/pended_on for threads)
rather than the dticks-cancel check, the in-band cancel mechanism is
dead code. Remove it:

  - z_abort_timeout(): no callers; remove the function definition
    in kernel/timeout.c and the declaration in kernel/include/timeout_q.h.
  - z_is_timeout_handler_canceled(): no callers (handlers no longer
    bail on dticks); remove.
  - TIMEOUT_DTICKS_ANNOUNCING: nothing reads it; remove the macro and
    the corresponding t->dticks = ANNOUNCING write in
    sys_clock_announce_locked()'s dispatch loop.

TIMEOUT_DTICKS_ABORTED stays. z_is_aborted_timeout() (used by
kernel/sleep.c via z_is_aborted_thread_timeout to distinguish a
k_wakeup'd thread from one that timed out) reads it, and
z_try_abort_timeout() still writes it on the linked-removal and
the same-CPU IRQ best-effort paths.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
59cf34bf21 kernel: work: migrate to z_try_abort_timeout()
Migrate work.c's two abort sites away from z_abort_timeout() and the
dticks-cancel mechanism.

Work queue timeout monitor (CONFIG_WORKQUEUE_WORK_TIMEOUT):

  work_timeout_stop_locked() uses a best-effort (void)
  z_try_abort_timeout() -- it does not wait for an in-flight handler.
  That abort flags the in-flight timeout superseded, and
  work_timeout_handler() bails on it via z_timeout_inflight_superseded(),
  in addition to the existing queue->finished check.

  Both checks are required. queue->finished alone is insufficient:
  if the work thread completes the item and starts a new one before
  this handler wins <lock>, finished is reset to false and the stale
  handler would abort the wrong (newly-started) worker. The superseded
  flag is what distinguishes "this timeout was aborted" from "a new
  item is being monitored". This restores, in the new mechanism, the
  cancellation guarantee that z_is_timeout_handler_canceled() provided
  on main.

Delayable work (work_timeout / unschedule_locked):

  work_timeout() drops the z_is_timeout_handler_canceled check and
  relies on K_WORK_DELAYED_BIT, the existing wake-ownership flag:
  whichever side -- this handler or unschedule_locked() -- atomically
  clears it via flag_test_and_clear() first owns the outcome; the
  other observes the cleared bit and bails. This is sufficient here
  (no superseded check needed) because unschedule_locked() waits:

  unschedule_locked() takes the caller's sched-lock key as a pointer
  and, on the success path, waits for any in-flight handler to fully
  complete before returning true. This closes a pre-existing UAF
  window: when unschedule_locked() returned true with the old
  z_abort_timeout(), the handler could still be about to run
  (blocked on work.c::lock while we held it), and would dereference
  work->flags after the caller released the lock. If a user-space
  caller freed dwork between k_work_cancel_delayable() returning
  and the handler bailing, the read raced with the free. The wait
  also serializes against k_work_schedule_for_queue() re-arming the
  same dwork for a new schedule -- without it, the old handler could
  unblock after the re-arm and erroneously submit against the new
  setup.

  cancel_delayable_async_locked() and the three callers of
  unschedule_locked() (cancel_delayable_async_locked,
  k_work_reschedule_for_queue, k_work_flush_delayable) forward the
  key pointer through, and the two callers of
  cancel_delayable_async_locked (k_work_cancel_delayable,
  k_work_cancel_delayable_sync) pass &key.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
c5b4762c36 kernel: poll: migrate to z_try_abort_timeout()
triggered_work_expiration_handler and signal_triggered_work /
triggered_work_cancel both want to take a k_work_poll item from
"waiting on event" to "submitted to work queue". Pre-1b8c7a3 the
dticks-cancel check in the handler made it bail when the cancel
side had aborted the timeout; without it, the in-flight handler
could submit the work twice and overwrite poll_result.

Use poller.is_polling as the wake-ownership flag. signal_poll_event
already clears it after a successful signal; the migration makes
signal_triggered_work() and triggered_work_cancel() clear it BEFORE
the abort, and adds a check at handler entry so the loser bails.
All under poll.c::lock, so the flip is atomic between the three
sides.

z_abort_timeout() is replaced by a best-effort (void)
z_try_abort_timeout() at the signal_triggered_work site -- a racing
in-flight handler observes is_polling=false and returns without
doing work.

triggered_work_cancel() additionally waits for any in-flight handler
to fully complete before returning success. This closes a pre-existing
UAF window: when cancel returned 0, the handler could still be about
to run (blocked on poll.c::lock while we held it), and would
dereference work->poller.is_polling after the caller released the
lock. If a user-space caller freed `work` between k_work_poll_cancel()
returning and the handler bailing, the read raced with the free. The
same wait also serializes against k_work_poll_submit_to_queue()
re-arming the same `work` for a new poll: without it, the old
handler could unblock after the re-arm flipped is_polling back to
true and erroneously act on the new setup.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00
Nicolas Pitre
f5dfc4ba3e kernel: timeslicing: migrate to z_try_abort_timeout()
z_time_slice_reset() calls z_abort_timeout() to cancel the previous
slice timeout before scheduling the next. The slice_timeout handler
only sets a per-CPU flag and flags an IPI -- it does not take
_sched_spinlock or dereference thread storage -- so a best-effort
(void) z_try_abort_timeout() is sufficient: a racing in-flight
handler sets slice_expired[cpu], which the trailing
"slice_expired[cpu] = false" assignment in z_time_slice_reset()
overwrites if our caller cares; if the handler runs after the clear,
the next dispatch picks up the (intentional) "expired" state and
rearms. No correctness concern either way.

Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
2026-06-15 18:35:49 -04:00