Replace the lock/test/restore dance used to probe the current IRQ state
with a direct non-modifying read:
- z_spin_is_locked() (UP path) simply negates
arch_cpu_irqs_are_enabled().
- k_can_yield() and z_smp_cpu_mobile() likewise drop their lock/unlock
pair.
- arch_spin_relax() asserts IRQs are disabled without the sneaky
unpaired arch_irq_lock() it used to rely on.
- tests/arch/arm/arm_no_multithreading: same simplification on a
probe-only assertion.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Extend z_spin_is_locked() to non-SMP configurations so assertions like
the one in z_unpend_all_locked() can validate lock ownership in UP
builds too. In UP a spinlock reduces to an IRQ lock, so the check
samples the current IRQ state via arch_irq_lock() / arch_irq_unlock().
Drop the now-unnecessary CONFIG_SMP guard around the sched spinlock
assertion in z_unpend_all_locked().
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
This moves the atomic_c.c from kernel to lib/os as atomic
functions are not exactly kernel features.
This also moves all the atomic kconfigs from kernel to lib/os
as the atomic headers are already under include/zephyr/sys/.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
errno is not exactly a kernel functionality but more of C
library feature. So move errno from kernel into lib/libc/common.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
k_free() bypasses k_heap_free() to avoid scheduler lock involvement,
going directly to sys_heap_free() instead. This means nothing in
kheap.c may have any callers, and since it is in a library linked
without --whole-archive, the linker may discard it entirely.
However, kheap.c contains a SYS_INIT handler that initializes all
statically defined k_heap objects (those created with K_HEAP_DEFINE).
Without it, heaps such as the system heap or those used as thread
resource pools via k_thread_heap_assign() are never initialized:
their internal sys_heap pointer remains NULL, causing a crash on
the first allocation.
This can be reproduced without this commit with e.g.:
west build -b qemu_cortex_a53 tests/kernel/poll
Force kheap.o into the link by adding a __used reference to
k_heap_init in mempool.c. This is enough to pull in kheap.o and
its SYS_INIT registration. Unused functions from kheap.o are
still garbage collected.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
k_free() now goes directly to sys_heap_free() under heap->lock,
bypassing k_heap_free() and its z_unpend_all() call. This is
symmetric with z_alloc_helper() which already bypasses k_heap_alloc()
to go directly to sys_heap_*().
This avoids any scheduler lock involvement in the k_free() path,
eliminating the recursive _sched_spinlock issue when k_free() is
called from halt_thread() during thread abort with CONFIG_USERSPACE
and CONFIG_DYNAMIC_OBJECTS.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
This partially reverts commit 9cef0da05c ("kernel: avoid recursive
scheduler lock in k_heap_free path"), keeping only the sched.c changes
(z_unpend_all_locked / z_unpend_all refactoring).
The _sched_locked variants of k_free, k_heap_free, k_msgq_cleanup,
k_stack_cleanup and the sched_locked parameter plumbing through
unref_check were an overcomplicated approach. A simpler fix follows.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
This moves boot arguments from kernel into the lib/os.
This is not strictly a kernel function so this change provides
a separation between core kernel functionalities and others.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
Boot banner is not exactly a kernel feature. It is more like
an OS feature so moving it into lib/os.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
IAR emits diagnostic Go004 ("function cannot be inlined") for
every ALWAYS_INLINE function when optimisation is disabled, e.g.
in debug builds. The previous workaround wrapped each affected
function in per-function preprocessor guard pairs:
#ifdef IAR_SUPPRESS_ALWAYS_INLINE_WARNING_FLAG
TOOLCHAIN_DISABLE_WARNING(TOOLCHAIN_WARNING_ALWAYS_INLINE)
#endif
static ALWAYS_INLINE void foo(...) { ... }
#ifdef IAR_SUPPRESS_ALWAYS_INLINE_WARNING_FLAG
TOOLCHAIN_ENABLE_WARNING(TOOLCHAIN_WARNING_ALWAYS_INLINE)
#endif
This pattern is highly intrusive, scatters toolchain-specific
knowledge across generic source files, and requires a guard pair
every time a new ALWAYS_INLINE function is added for IAR.
Replace it with a single override of ALWAYS_INLINE inside
iccarm.h, using the C99 _Pragma operator to embed the diagnostic
suppression in the macro itself.
Assisted-by: GitHub Copilot:claude-sonnet-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Export the resolved K_HEAP_MEM_POOL_SIZE value to the
linker generator and evaluate it while generating the
IAR command file.
Fixes: #107234
Signed-off-by: Jason Yu <zejiang.yu@nxp.com>
gen_offset.h is an architecture-specific header, not a kernel one.
Move it under the arch tree where it belongs.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Migrate scheduler API implementations (k_sched_lock/unlock,
z_reschedule, z_yield_current, etc.) and their private declarations
from ksched.h/sched.c into scheduler.c and scheduler.h.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Move time-slice related declarations from ksched.h into the
dedicated kernel/include/timeslicing.h header.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Migrate z_add_thread_to_ready_q(), z_remove_thread_from_ready_q(),
and related helpers from sched.c to scheduler.c.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Move z_sched_init to scheduler.c and somplify implementation getting rid
of single use init_ready_q.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Reorder so that z_ready_thread and z_unready_thread are adjacent,
improving code locality for related queue operations.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Move run-queue management functions (add/remove/peek thread,
choose_next_thread) from sched.c into the new
kernel/include/run_q.h header.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Move meta-IRQ (highest-priority cooperative queue) scheduling
functions from sched.c into a new kernel/include/metairq.h header
to reduce sched.c size and group related logic.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Reduce complexity of sched.c by encapsulating sleep handling code
(k_sleep, k_usleep, k_msleep) into its own file.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Relocate k_thread_start(), k_thread_abort(), k_thread_suspend(), and
k_thread_resume() from sched.c to thread.c alongside related thread
lifecycle code.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
SYS_PORT_TRACING_OBJ_FUNC_EXIT fired while the spinlock was still
held. The standard pattern across all other heap/kernel-object
functions is to release the lock first, then emit the EXIT trace.
Swap the two lines so the unlock precedes the tracing call.
Assisted-by: GitHub Copilot:claude-opus-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
In z_impl_k_stack_pop, SYS_PORT_TRACING_OBJ_FUNC_BLOCKING was emitted
before the K_NO_WAIT timeout check. When the stack is empty and
timeout == K_NO_WAIT, the function emitted the BLOCKING trace and then
immediately returned -EBUSY without ever blocking. Tracing consumers
that expect a thread block to follow each BLOCKING event would observe
a spurious BLOCKING with no corresponding suspend.
Move the BLOCKING trace to after the K_NO_WAIT early-return so it is
only emitted when the thread is actually about to pend.
Assisted-by: GitHub Copilot:claude-opus-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
In mbox_message_put, when an async (dummy thread) sender matches a
waiting receiver, the function calls z_reschedule() and returns 0
without emitting SYS_PORT_TRACING_OBJ_FUNC_EXIT. Every other return
path from mbox_message_put emits the EXIT trace before returning.
This missing trace leaves a dangling ENTER event for tracing consumers
(e.g. Percepio TraceRecorder) that expect matched ENTER/EXIT pairs.
Add the missing SYS_PORT_TRACING_OBJ_FUNC_EXIT call after
z_reschedule() on the async early-return path, matching the result
value 0 used by all other successful paths.
Assisted-by: GitHub Copilot:claude-opus-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
queue_insert emitted SYS_PORT_TRACING_OBJ_FUNC_BLOCKING twice:
1. When a pending thread is found and woken (correct).
2. Unconditionally before sys_sflist_insert when no thread is
pending and the item is placed directly on the list (incorrect).
The second emission is wrong: no blocking occurs in that path — the
caller's data is simply enqueued and the function returns. Emitting
a BLOCKING event there misrepresents the operation to tracing
consumers and is likely a copy-paste error from the first branch.
Remove the second SYS_PORT_TRACING_OBJ_FUNC_BLOCKING call.
Assisted-by: GitHub Copilot:claude-opus-4.6
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
This is can be used to revoke access from all but the current
thread, which is useful when reassigning an object without having
to worry about previous permissions.
Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
This fixes a subtle race-condition in the k_timer expiration
handler z_timer_expiration_handler(). There was a small window
of opportunity between when sys_clock_announce() unlocked
interrupts and that handler re-locked them that one or more
higher priority interrupts (or threads running on another CPU
if in an SMP environment) could not only abort the ktimer's
timeout, but restart it as well. Both of these situations are
now detectable in the handler (resulting in an immediate return
from the handler).
To make this work, every case where the ktimer internals either
adds or aborts its timeout is now encapsulated by the ktimer lock.
Thus, when the handler tests if the timeout handler has been
canceled with only the ktimer lock being held, we know that no
other thread or ISR can be modifying the ktimer's timeout.
Fixes#106654
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
On SMP systems with tickless kernels, a race condition exists between
timer driver ISRs and the kernel's tick accounting. The driver updates
its hardware cycle baseline under a private lock, then calls
sys_clock_announce() which updates curr_tick under the separate
timeout_lock. In the gap between these two lock releases, any kernel
code calling sys_clock_elapsed() sees the new driver baseline but the
old curr_tick, producing inconsistent time values that can go backwards.
This affects every code path using the internal elapsed() helper:
uptime queries, timeout scheduling, timeout cancellation, remaining
time queries, and next-expiry calculations.
The root cause is two separate locks protecting state that must be
mutually consistent. Fix this by exposing the kernel's timeout_lock
to timer drivers via sys_clock_lock()/sys_clock_unlock(), and
providing sys_clock_announce_locked() which assumes the lock is
already held.
Timer drivers can now acquire the single lock, update their hardware
state, and announce ticks all under the same lock — eliminating the
race window entirely. The key is passed to sys_clock_announce_locked()
which consumes it (releasing the lock when it returns).
The existing sys_clock_announce() becomes a backward-compatible wrapper,
allowing incremental driver migration with no flag day.
Document that sys_clock_set_timeout(), sys_clock_elapsed(), and
sys_clock_idle_exit() are called by the kernel with the timer lock
held. Update the timer driver guide in clocks.rst accordingly.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Add a runtime assertion in z_unpend_all_locked() to verify that
_sched_spinlock is actually held by the caller. This catches misuse
early given the function call depth involved.
Extend the availability of z_spin_is_locked() from CONFIG_SMP &&
CONFIG_TEST to also include CONFIG_ASSERT, so the check can be
used in __ASSERT() outside of test builds.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
When halt_thread() calls k_thread_perms_all_clear() under
_sched_spinlock, the permission cleanup can trigger k_free() on
dynamic objects. k_heap_free() then calls z_unpend_all() which
attempts to take _sched_spinlock again, causing a recursive lock.
Fix this by introducing k_heap_free_sched_locked() and
k_free_sched_locked() variants that use z_unpend_all_locked()
to operate on the wait queue without re-acquiring the scheduler
lock. The existing z_unpend_all() becomes a wrapper that takes
the lock and delegates to z_unpend_all_locked().
unref_check() gains a sched_locked parameter: the abort path
(clear_perms_cb) passes true to use the locked free variant,
while k_thread_perms_clear() passes false for the normal path.
Fixes#106659
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Move signal_pending_ipi() inside the K_SPINLOCK block in
z_get_next_switch_handle(). Calling it after the lock release creates a
window where a CPU can consume its own pending IPI bit via atomic_clear
in signal_pending_ipi(), then silently drop it in
arch_sched_directed_ipi() which skips the calling CPU (i == id).
In configurations where secondary CPUs have a single pinned thread and
take no timer or external interrupts, this can lead to a permanent hang:
the idle CPU can only be woken by IPIs, but no IPIs are pending and no
timeslicing IPIs will be generated since the idle thread is not sliceable.
This was reproduced when running under QEMU with the following sequence
of events observed:
CPU 0 CPU 1
───── ─────
Thread calls k_poll(K_MSEC(1))
z_pend_curr():
mark thread PENDING
z_add_timeout(1ms)
do_swap() to idle thread
WFI
Timer tick fires
sys_clock_announce():
slice_timeout(cpu1):
flag_ipi(BIT(1))
signal_pending_ipi():
MSIP[cpu1] = 1
CPU1 wakes from WFI
z_get_next_switch_handle():
acquire _sched_spinlock
next_up() → idle
(thread still PENDING,
timeout hasn't fired yet)
release _sched_spinlock
Timer tick fires
sys_clock_announce():
z_thread_timeout(thread):
z_unpend_thread(thread)
z_ready_thread(thread):
flag_ipi(BIT(1))
signal_pending_ipi():
atomic_clear(pending_ipi)
returns BIT(1)
arch_sched_directed_ipi(BIT(1))
skips self, IPI silently lost
return to idle thread
WFI
thread still on ready queue
Such an interleaving of events is, of course, likely only reproducible in
practice in virtualized environments where (v)CPUs can be descheduled.
With signal_pending_ipi() inside the lock, next_up() and the IPI
dispatch are atomic. Either the concurrent flag_ipi lands before the
lock is acquired (and next_up sees the thread), or it lands after the
lock is released (and the caller dispatches the IPI). There is no
window where a CPU can consume its own bit for a thread it hasn't seen.
Similar races exist in reschedule() and z_reschedule_irqlock() as well.
Although they won't cause the same permanent hang described above, it
can result in unnecessary rescheduling latency. Fix reschedule(), and
add a TODO to z_reschedule_irqlock(); it doesn't not currently take
the sched spinlock.
Signed-off-by: Andrew Bresticker <abrestic@meta.com>
Workq optionally yield after every work handler to avoid starving
other threads.
When current workq is empty after this work handler, current thread
will go to sleep in next loop. So no need to yield, bringing one more
schedule cost.
Signed-off-by: Fengming Ye <frank.ye@nxp.com>
Between the points in time when sys_clock_announce() calls the
timeout handler for delayable work and when that handler wins
the work queue spinlock another thread or ISR could have called
k_work_reschedule_for_queue(). Should this occur, the timeout
that the handler is trying to process becomes stale and the
handler should not proceed any further with it.
As the workqueue spinlock is the controlling lock (it is always
held before either aborting or adding a timeout), it is safe
for the handler to call z_is_timeout_handler_canceled() once
it holds the workqueue spinlock.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The workqueue work timeout feature is supposed to abort the work
queue thread if the time to execute a work item exceeds the work
queue's configured threshold. The work thread may race against the
timeout handler responsible for aborting the thread when the two
are running on separate CPUs--particularly since the timeout handler
only locks the workqueue spinlock for part of its duration.
To get around this, two separate flags must be checked a 'finished'
flag to indicate that the thread has finished processing the work
item and the timeout's flag indicating if it has been removed while
processing its timeout handler. Should either be found to be true
within in the timeout handler, the thread is deemed to have completed
in time and the timeout handler proceeds no further.
Otherwise the timeout handler is deemed to have won the race and the
workqueue thread is aborted. Should the workqueue thread detect this,
it goes to sleep until it can be aborted to prevent it from handling
any more work items.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>