printk_unlocked() serves callers that know they cannot lock, not the ones
that find out too late. A fault taken inside printk() leaves the spinlock
held by a context about to die, so every later printk() blocks and the
crash report is lost along with everything after it. CONFIG_SPIN_VALIDATE
makes it louder rather than better: validation fails, the assertion is
reported through printk, and the system spins there repeating itself.
That happens on a uniprocessor too, where the lock is uncontended but
still recorded as held.
Converting call sites cannot fix it, because code reached after the
failure does not know it is in a crash. So make printk() itself aware:
printk_panic() switches it to the unlocked path and the reporting paths
call it before emitting anything.
The switch is one way. Clearing it would also require releasing the
orphaned lock, and telling "orphaned" from "held by a CPU that is still
running" needs the owner, which only CONFIG_SPIN_VALIDATE records. So a
system that has reported a fatal error keeps unserialized output for the
rest of its life, which costs interleaving and no content.
Linux does the same: bust_spinlocks(), called from its oops and panic
paths, raises oops_in_progress so printk skips console locking. This
mechanism is narrower, printk simply stops locking, but the reasoning is
identical.
Because it cannot be undone, it must only be thrown once the fault is
terminal, which rules out the architecture entry points: several return.
z_arm64_fatal_error() resumes once demand paging has serviced the fault,
so switching there would cost a paging system its printk locking on the
first page fault. EXCEPTION_DUMP() runs only after the recoverable cases
are ruled out and also covers arch-specific dumps such as the Cortex-M
one behind PR_FAULT_INFO, and z_fatal_error() covers reports with no
dump, like k_panic() from a failed assertion. OpenRISC switches at entry
because it reports with LOG_ERR(), which is safe there since its handler
never returns.
k_str_out() shares the lock and honours the switch, so crashes reported
through printf() or puts() are covered too.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
The internal kernel routine mbox_message_dispose() was non-atomically
modifying the internal thread fields 'thread_state'. To correct this,
_sched_spinlock must be held to prevent corruption of this field from
either an ISR or another CPU acting upon the thread.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The mailbox code must search the waiting threads for a matching
message. Prior to this, it was doing so the using _WAIT_Q_FOR_EACH()
macro while holding the mailbox' spinlock. That pattern, however,
does not provide sufficient protection on an SMP system as another
CPU could process a timer that expires one of those waiting threads
thereby unexpectedly mutating the wait queue during the search. This
in turn could cause corruption and a crash.
To work around this, z_sched_waitq_walk() must be used to iterate
through the wait queue to search for and subsequently act upon
a matching message.
Note: z_sched_waitq_walk() locks _sched_spinlock for the duration of
the search. Consequently, if a thread expiration were to occur during
the search, it would spin in z_thread_timeout() until it wins
_sched_spinlock (after the search is complete).
Fixes#111643
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The mailbox code no longer reaches deep into the thread structure
to determine if a thread is a dummy thread. The kernel has an
existing helper routine called 'is_thread_dummy()'. Use that instead.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
Updates the z_sched_waitq_walk() documentation to indicate that the
the walk_func callback may safefly remove the thread identified by
the callback's argument from the wait queue on the final iteration.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The existing implementation unconditionally restores owner_orig_prio on
unlock, ignoring other held mutexes that still have high-priority waiters.
This causes incorrect priority restoration in nested mutex scenarios and
requires mutexes to be released in strict reverse acquisition order.
Additionally, priority boosts were not propagated through ownership chains
when the mutex owner was itself blocked on another mutex.
This change adds per-thread held_mutexes tracking and mutex_pended_on
pointer to enable:
- Correct priority recalculation on unlock by scanning all remaining held
mutexes
- Chained priority inheritance through the full ownership chain
- Deadlock detection: assert on K_FOREVER circular ownership where
every cycle member also waits forever; bounded chain walk prevents
livelock on cycles not involving the current thread
- Per-thread orig_prio field records true pre-inheritance priority,
fixing priority floor when mutexes are released in non-LIFO order
struct k_thread grows by three pointers plus a byte in default builds
(held_mutexes, mutex_pended_on, orig_prio).
Signed-off-by: Mayur Salve <msalve@qti.qualcomm.com>
Calling thread_schedule_new with a timeout not equal to K_NO_WAIT will
result in a z_add_timeout call, in the much common configuration
CONFIG_SYS_CLOCK_EXISTS=y, which returns early on K_FOREVER making the
compare redundant.
When !CONFIG_SYS_CLOCK_EXISTS the behavior is kept such that a timeout
value of K_FOREVER still results in a nop.
Signed-off-by: Emil Hammarström <emil.a.hammarstrom@gmail.com>
MISRA C:2012 Rule 8.2 requires every parameter in a function type to be
named. The local prototype for main() used when CONFIG_BOOTARGS is
enabled declared its two parameters with bare types.
Declaration only, no functional change.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
MISRA C:2012 Rule 8.2 requires every parameter in a function type to be
named, and that includes definitions, not just declarations. The two
object core walk functions were declared with named callback parameters
but still defined with unnamed ones.
Use the same names as the header so declaration and definition agree.
No functional change.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
The non-SMP branch only asserted that cpu is 0 and silently
returned CPU 0 statistics for any out-of-range value with
CONFIG_ASSERT=n. The SMP branch had no validation at all and
passed the value through to z_sched_cpu_usage(), indexing the
per-CPU array out of bounds.
Validate cpu against arch_num_cpus() with CHECKIF() returning
-EINVAL, consistent with the existing NULL check for stats, and
unify the SMP and non-SMP branches since arch_num_cpus() is 1 in
the latter case.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Several demand paging failure conditions were only caught by
assertions, although they are reachable at runtime in a correct
program and were silently ignored with CONFIG_ASSERT=n:
- k_mem_paging_eviction_select() returning NULL (all evictable
page frames pinned or busy) in map_anon_page() and
do_page_fault() led to a NULL dereference. Return -ENOMEM from
map_anon_page() and fail the fault in do_page_fault().
- page_frame_prepare_locked() failing with -ENOMEM (backing store
full) in do_page_fault() continued with an unprepared page frame
and garbage page-out location. Fail the fault instead, matching
the handling map_anon_page() and do_mem_evict() already have.
- do_page_in()/do_mem_pin() ignored a false return from
do_page_fault() (unmapped address, and now also paging resource
exhaustion). k_mem_page_in()/k_mem_pin() return void, so the
caller would continue believing memory is resident/pinned and
the failure would surface at an arbitrary later access, possibly
from a context that cannot handle a page fault. Log an error and
panic instead.
Also drop a tautological assertion in do_mem_evict() that was
only reachable when its condition was already true.
Failed page faults are reported as fatal errors by the arch fault
handler, as already done for ARCH_PAGE_LOCATION_BAD.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Disabling stats on the current thread accounts the in-progress window
into the CPU stats but forgets to update cpu->usage0. The next
z_sched_usage_stop() recomputes the same window and adds it again.
Signed-off-by: Yiren Guo <guoyr_2013@hotmail.com>
If logging is disabled, we want to get the same exception and
information, right now half of the information is being dropped as they
go through LOG_ERR only while the rest go via EXCEPTION_DUMP/printk.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
k_sleep() and k_usleep() resolve to an out of line implementation living in
another translation unit, a system call under CONFIG_USERSPACE and a plain
call otherwise. Either way the compiler sees neither whether the requested
duration is a constant nor whether the caller cares about the returned
value, so it always emits the tick to millisecond and microsecond
conversions. Those conversions are the only thing that drags the 64 bit
division helper into many small builds: CONFIG_SYS_CLOCK_TICKS_PER_SEC
defaults to 10000 on tickless platforms, which lands in the integer
division path and divides a 64 bit value by ten, something GCC will not
strength reduce on a 32 bit target.
Promote the former z_tick_sleep() helper to a k_sleep_ticks() system call
that reports the remaining time in ticks, and rebuild k_sleep() and
k_usleep() as inlines on top of it. The conversions now live at the call
site where the compiler can fold or discard them. Nearly every caller
discards them: of the 4036 sleep call sites in the tree, exactly six use
the returned value, and only three of those are outside of tests.
What this removes is best seen in the two out of line entry points that
disappear, which on a nucleo_f030r8 (Cortex-M0, 10000 ticks/s) were the
only two callers of the 64 bit division helper in the whole image:
08002594 <z_impl_k_sleep>:
bl 8001... <z_tick_sleep>
...
movs r0, #9 /* the ceil() bias, 10 - 1 */
adds r0, r0, r2
adcs r1, r3
movs r2, #10 /* ticks / 10 -> milliseconds */
bl 8000180 <__aeabi_uldivmod>
080025c0 <z_impl_k_usleep>:
movs r0, #99 /* the ceil() bias, 100 - 1 */
...
movs r2, #100 /* microseconds / 100 -> ticks */
bl 8000180 <__aeabi_uldivmod>
bl 8001... <z_tick_sleep>
Both are gone, and so are __aeabi_uldivmod and __udivmoddi4, which are no
longer referenced anywhere. A k_usleep(250) call site now passes the
literal 3 ticks the compiler worked out on its own, where it used to pass
250 microseconds to a helper that divided by 100 at runtime. For a loop
calling k_msleep(100), k_sleep(K_SECONDS(1)) and k_usleep(250) without
using the results, FLASH drops from 11028 to 10600 bytes.
Tracing stays in the out of line implementation. It cannot move into the
inlines because every tracing backend header includes kernel.h, so a
translation unit that reaches a backend header first would expand
SYS_PORT_TRACING_* before those macros exist. A consequence is that the
sleep_exit hook now reports ticks rather than milliseconds and the usleep
hook is no longer emitted; both are covered in a following commit.
The microsecond duration is clamped to zero the way Z_TIMEOUT_US() already
does. k_usleep() never did that, but it used to truncate the 64 bit
conversion into an int32_t before handing it over, which hid the worst of
it.
Passing the conversion through whole would turn a negative duration into a
sleep of roughly 1.8e17 ticks rather than a bogus absolute deadline.
This also removes a potential link error. k_usleep() had no
implementation at all with CONFIG_MULTITHREADING=n, since nothread.c only
ever provided z_impl_k_sleep(), so any caller failed to link. Nothing in
the tree happens to call it in that configuration, which is presumably why
it went unnoticed; a single z_impl_k_sleep_ticks() now serves both.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
k_sleep(), k_msleep() and k_usleep() are declared in kernel.h today.
Later commits grow that API with a tick based primitive plus inline unit
conversions, which is more than kernel.h ought to carry for a single
service.
Move the three declarations verbatim into a new include/zephyr/sleep.h
and have kernel.h include it, so existing users need no change. The new
header has to be registered with zephyr_syscall_header(), otherwise the
marshalling stubs for k_sleep() and k_usleep() stop being generated and
CONFIG_USERSPACE builds fail to find k_sleep_mrsh.c.
No functional change. A test application built for nucleo_f030r8 comes
out byte identical, at 11028 bytes of FLASH before and after.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
The idle path asks this for the time until the next wakeup, and gets a
tick count that never says "there is nothing to wake up for": an empty
timeout list is reported as the capped announce budget, exactly like a
deadline further out than can be programmed in one step. The power
management code cannot then tell the two apart, so it arms the timer in
both cases and a system with sloppy idle enabled keeps waking up for
nothing.
Make the same decision reprogram_next() makes. With
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE allowing the uptime to drift, an empty
list is reported as K_TICKS_FOREVER; anything else is a wait, whether a
real deadline or the synthetic one that keeps the announce range
covered. Without sloppy idle the empty case stays a wait, so the timer
remains armed and timekeeping is unaffected.
The return type becomes unsigned to match the tick type used throughout
the timer interface. The conversion at the only caller, in the idle
path, is value preserving in both directions: every wait is capped at
SYS_CLOCK_MAX_WAIT, which is INT32_MAX, and K_TICKS_FOREVER is the same
value read either way. Nothing outside the kernel sees the change; the
power management API keeps its signed tick counts.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
An empty timeout list normally still gets the driver a wakeup request.
That is a synthetic timeout: nothing is waiting for it, it is there to
keep the announce baseline moving so uptime stays correct.
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE is the permission to skip it, which is
why an empty list only becomes visible to a driver when that option is
set, and today it becomes visible as a tick value the driver has to
recognise rather than as a condition of its own.
Add a weak sys_clock_no_timeout() hook, called in place of
sys_clock_set_timeout() when the list is empty and sloppy idle allows the
uptime to drift. Its default asks sys_clock_set_timeout() for UINT32_MAX
ticks, which is what K_TICKS_FOREVER is in the unsigned tick domain this
interface now uses, so an out-of-tree driver that has not migrated still
sees the "no deadline" value it always did. What gets deprecated is the
special meaning, not the call.
next_timeout() no longer decides sloppy idle, so its empty-list and
far-timeout arms collapse into one capped budget. The far-timeout arm no
longer stops a tick short of the cap either: it only did so to stay
distinguishable from the empty case, which now travels out of band.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
In commit 4f9df472fc1d06a6deeb7dddf6ab167c7cba984f
TOOLCHAIN_DISABLE_WARNING with a hard coded argument in work.c. This
created warnings when using the iar toolchain. Fix it by using a
preprocessor macro.
Signed-off-by: Gustav Holmberg <gustav.holmberg@qt.io>
kernel_offsets.h pulls in three headers before defining its own
include guard, so the guard is not the first thing in the file as
MISRA C:2012 Directive 4.10 requires.
Move the guard above the includes. The header is still only used by
the offsets generator, and the includes stay inside the guard.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
'exit' is reserved by the C standard for any use, so using it as a
label violates MISRA C:2012 Rule 21.2. Rename the labels in the
message queue and pipe implementations to 'out', which is the spelling
already used elsewhere in the kernel.
Pure rename, no functional change.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
CONFIG_SMP_BOOT_DELAY was a global, application-level switch that made
the kernel skip starting every secondary CPU at boot, used in-tree only
by two tests and in practice by one platform class (intel_adsp, where
the host or PM policy brings DSP cores up on demand). Boot topology is
a hardware/platform property, and all-or-nothing is needlessly coarse.
Replace it with a per-CPU devicetree flag, zephyr,deferred-start, on
the /cpus children (mirroring zephyr,deferred-init for devices):
z_smp_init() now always runs and simply skips flagged CPUs, which are
brought up at run time with the existing k_smp_cpu_start() (or
k_smp_cpu_resume()). Deferral is per CPU, so asymmetric bring-up
(start some cores at boot, defer others) is now expressible, and the
special-case branch disappears from the boot path.
The flag is declared in the common cpu.yaml binding. The lookup uses
DT_PROP_OR() so cpu nodes whose binding does not cover the property
simply cannot be deferred rather than breaking the build; a binding
for the intel,x86_64 compatible used by qemu_x86_64's cpu nodes was
missing entirely and is added.
The two users are converted: tests/kernel/multiprocessing/
smp_boot_delay marks the secondary CPUs in per-board overlays (both
tests pass on qemu_x86_64) and tests/boards/intel_adsp/smoke gains
overlays for its four platforms (all build). Normal SMP boot is
unaffected (verified on qemu_x86_64, all CPUs online).
Out-of-tree users migrate by dropping CONFIG_SMP_BOOT_DELAY=y and
adding zephyr,deferred-start to the deferred cpu nodes in their board
overlay.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Clang (pre-v20) enables frame pointers at all optimization levels,
unlike gcc which only enables them only at -O0. This causes the
same r7 clobber issue in arch_switch inline assembly as for GCC,
if CONFIG_SWITCH=y.
Extend the -fomit-frame-pointer workaround to apply unconditionally
for clang and also add sleep.c and thread.c which also started to
need the flag since the workaround was added.
Signed-off-by: Tim Pambor <tim.pambor@codewrights.de>
z_vrfy_device_deinit() validated its argument with K_OBJ_ANY, so any kernel
object the calling thread has access to could be passed as a struct device.
z_impl_device_deinit() then calls dev->ops.deinit(dev), an indirect call
through a function pointer read out of the confused object, giving user
mode an arbitrary call primitive when CONFIG_DEVICE_DEINIT_SUPPORT=y.
Use K_OBJ_DRIVER_ANY, matching z_vrfy_device_init() and
z_vrfy_device_is_ready().
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
dynamic_object_create() computed the allocation size as
obj_size_get(otype) + size, where size comes straight from user mode via
the k_object_alloc_size() syscall - the verifier is a bare pass-through and
z_object_alloc() only range-checks otype. A size close to SIZE_MAX wraps,
so the allocation is small while the object is tagged with the requested
type; the matching init syscall then writes a full object over it. The
K_OBJ_THREAD_STACK_ELEMENT branch has the same problem through
STACK_ELEMENT_DATA_SIZE(), which rounds up and adds overhead.
Reject both overflows and free the descriptor.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Updates k_work_submit() to use z_work_submit_to_queue() instead of
k_work_submit_to_queue() to remove superfluous tracing calls.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
Migrates the schedule point from k_work_submit_to_queue() to the
internal routine z_work_submit_to_queue().
Though the function description previously indicated that
z_work_submit_to_queue() was not to yield the thread, investigation
showed that its only use outside the work.c file was poll.c and that
interrupts are locked at each of those call points. As interrupts
were already locked prior to calling this routine from poll.c,
z_reschedule() will not swap out the current thread and the behavior
remains unchanged.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
z_add_timeout returns without any effect if timeout is equal to K_FOREVER
so there is no need to guard its timeout parameter from a K_FOREVER.
Signed-off-by: Emil Hammarström <emil.a.hammarstrom@gmail.com>
There are a few functions inside kernel/sched.c that are simply
one line wrappers to functions having the exact same signatures.
Unifying the naming would be ideal but the underlying functions
are being called by others under different kernel namespace.
So for now, the simplest is to use function aliasing and let
linker do its magic to avoid unnecessary trampoline. Saves
a few CPU cycles too.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
With CONFIG_SPIN_VALIDATE, z_assert_can_swap() & halt_thread() check if
the value set in base.swap_data corresponds to &z_spinlock_abort_sentinel
(it will be set to this value in
subsys/testsuite/ztest/src/ztest_error_hook.c if there is a fatal error
or assert).
But nobody is initializing this (struct k_thread).base.swap_data to
anything.
So when there is no failures (or no ztest code), we are checking a random
value and potentially doing weird things.
Let's initialize it to NULL to avoid this.
This check was added in d844a4a861
Signed-off-by: Alberto Escolar Piedras <alberto.escolar.piedras@nordicsemi.no>
Moves the assert checking that exactly one CPU in the mask is set
to be inside the locked section. This closes a hole where another
CPU could have called cpu_mask_mod() and performed an OR operation
setting more bits just before the current CPU read the cpu_mask
field leading to an incorrect assert evaulation.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
Header file include/zephyr/sys_clock.h is deprecated and will be removed
someday. Update the whole file tree to include zephyr/sys/clock.h
straight instead of zephyr/sys_clock.h.
This change was made running the sed shell command below:
$ sed -i 's/zephyr\/sys_clock\.h/zephyr\/sys\/clock\.h/' \
`grep -rsl "zephyr/sys_clock\.h" kernel/`
Signed-off-by: Etienne Carriere <etienne.carriere@st.com>
A blocking pend from ISR context is a programming error with no safe
recovery: it attempts to sleep whatever thread was interrupted and, on
CONFIG_SWAP_NONATOMIC, corrupts the scheduler (the qnode_dlist
double-link / NULL-deref traced in #111518).
Catch it at the funnel instead of per-API: every blocking primitive
(sem, mutex, msgq, queue, stack, pipe, poll, events, mailbox, mem_slab,
kheap, futex, ...) routes through z_pend_curr, and reaching it always
means a real block (callers handle K_NO_WAIT beforehand). So one check
there covers them all, with no per-API return-contract churn:
if (arch_is_in_isr()) {
__ASSERT(false, "blocking pend from ISR context");
k_panic();
}
__ASSERT(false, ...) gives the message and backtrace in debug builds;
k_panic() makes it fatal in every build, closing the production gap
(CONFIG_ASSERT defaults to n) that let the misuse ship and run for
hours. Located before _sched_spinlock is acquired, so no lock is held
on the panic path.
Also assert in add_to_waitq_locked() that a thread is not already on a
wait queue when added to a new one (pended_on == NULL) -- belt-and-
suspenders for a double-pend that reaches the scheduler despite the
above, e.g. a non-current thread re-pended without an intervening
unpend.
Suggested-by: Nicolas Pitre <npitre@baylibre.com>
Signed-off-by: Tibor Kiss <kiss.tibor@gmail.com>
z_add_timeout returns without any effect if timeout is equal to K_FOREVER,
and thus z_add_thread_timeout does this by extension and there is no need
to guard its timeout parameter from a K_FOREVER.
Signed-off-by: Emil Hammarström <emil.a.hammarstrom@gmail.com>
Move the declaration of arch_page_phys_get() from the internal
kernel_arch_interface.h to the public arch_interface.h, under the
existing arch-mmu group, so that drivers and subsystems can query the
physical address of an already-mapped virtual page without depending
on private kernel headers.
The declaration stays unconditional, as it was in the internal header:
callers such as munmap() in subsys/portability/posix/options/mmap.c
reference the function even when CONFIG_MMU is disabled (guarded only
by a runtime IS_ENABLED() check), so hiding the declaration behind
CONFIG_MMU breaks the build on MMU-less platforms.
Clean up the existing users that resorted to including the internal
header or to extern declarations:
- subsys/portability/posix/options/mmap.c
- subsys/portability/posix/options/shm.c
- subsys/portability/cmsis_rtos_v1/cmsis_thread.c
Fixes#113314
Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
There is no need to grab CPU ID from _current_cpu->id down in
the ipi_work_process() call chain, as it can be passed as
a function argument and can be reused. Using _current_cpu can
be costly with assertion on as it calls z_smp_cpu_mobile for
each call.
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
z_main_thread is only needed if multithreading is enabled, so move its
definition inside a multithreading check.
Signed-off-by: Josh DeWitt <josh.dewitt@garmin.com>
With low-power idle handoff now going through sys_clock_idle_enter(),
nothing passes idle == true to sys_clock_set_timeout() any more, so the
argument is dead. Drop it: sys_clock_set_timeout(uint32_t ticks) now
means exactly "program the next tick, N ticks out", nothing else.
This updates the prototype, the weak default, every in-tree timer driver
definition (including the out-of-tree-style board timer under boards/),
and the core call sites. For the five drivers with low-power behaviour
this only removes the now-unused idle argument that the previous change
left on set_timeout(); the handling itself stays in their
sys_clock_idle_enter().
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
The "no timeout pending, stop the clock" decision under
CONFIG_SYSTEM_CLOCK_SLOPPY_IDLE was expressed by next_timeout()
returning SYS_CLOCK_MAX_WAIT verbatim as a magic sentinel, which every
timer driver then had to recognise. Move the decision into the core and
give it an explicit, resumable interface.
Add a weak sys_clock_unused() hook (no-op default): the kernel calls it
when the timeout list is empty and sloppy idle allows uptime to drift,
in place of programming a wait. A driver may override it to actively
halt its counter; one that does not simply stops being reprogrammed and
quiesces on its own, so sloppy idle now works for every driver. Resume
is the next sys_clock_set_timeout(), exactly as before; the core keeps
no paused state.
next_timeout() no longer special-cases sloppy idle, so its empty-list
and far-timeout arms collapse to the same capped budget and
SYS_CLOCK_MAX_WAIT loses its sentinel meaning. The decision lives in
reprogram_next(), used only at the two sites where the list can drain
(abort, end of announce); the add path always has a pending timeout and
calls sys_clock_set_timeout() directly.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
The per-CPU loops in k_thread_runtime_stats_all_get() and the runtime
stats enable/disable paths used a uint8_t index while the loop bound
arch_num_cpus() is unsigned int.
Widen the index to unsigned int to match the bound type.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
The guard page checks in k_mem_unmap_phys_guard() asserted
ret == 0 inside an if (ret == 0) branch, so the assertions
could never fire and the "cannot find guard page" diagnostic
was unreachable in any build.
Assert the intended invariant (ret != 0, i.e. the guard page
is unmapped) before the branch, matching the assert-then-
handle pattern used elsewhere in this function. Behavior with
CONFIG_ASSERT=n is unchanged.
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
A CPU's usage counters only advance when it context switches, or when it
reports on itself through z_sched_cpu_usage(). Reading the runtime stats
of a *different* CPU therefore misses whatever time that CPU has spent in
its current thread since then.
The effect is worst for an idle CPU: a CPU that ran briefly and then went
idle has the active cycles of that run in its counters, but none of the
idle time that followed, because the idle thread has not been switched
out yet. Both k_thread_runtime_stats_cpu_get() and everything built on it
therefore see a window made up purely of execution cycles, i.e. a 100%
load for a CPU that is doing nothing.
Fold the in-progress cycles into the returned view, attributing them to
the idle thread or to the CPU depending on what it is currently running.
This is done for the returned stats only, without mutating the other
CPU's accounting, as that CPU will account for the same cycles itself
once it next updates. The read is consistent because usage_lock, which is
held here, also covers the usage0 timestamp and the counters.
Reading the current CPU is unchanged.
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
Several documentation defects in the two canonical arch interface
headers:
- The arch-timing group was defined twice, once in arch_interface.h
and once in kernel_arch_interface.h; turn the latter into an
addtogroup of the former.
- The arch-smp group was opened with addtogroup before its defgroup
appeared 300 lines later in the same file, leaving the SMP APIs
split across two blocks with the group defined mid-file; define
the group at first use and addtogroup at the second block.
- ARCH_STACK_PTR_ALIGN was the only stack macro documented with a
bare @def block and no __DOXYGEN__ stub definition, making it an
orphan @def; add the stub like its siblings.
- The ARCH_PCIE_IRQ_CONNECT doc block was wrapped in #ifdef
CONFIG_PCIE with no __DOXYGEN__ stub, so it never appeared in doc
builds; use the same __DOXYGEN__ pattern as the other macro slots.
- The arch-stackwalk group was closed with a malformed comment
block; use the standard group close.
- Fix a #return that should be @return in the
arch_thread_priv_stack_space_get() documentation.
In addition, give the ARCH_* macro slots real documentation: most of
them (ARCH_THREAD_STACK_RESERVED, ARCH_IRQ_CONNECT,
ARCH_PCIE_IRQ_CONNECT, ARCH_IRQ_DIRECT_CONNECT, the
ARCH_ISR_DIRECT_* family) carried nothing but a @see reference to
their public counterpart, leaving the porting-facing contract
undocumented. Describe what each hook must do and document its
parameters; add @brief lines to the stack alignment macros.
No functional change, comments and preprocessor-invisible
documentation stubs only.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
signal_pending_ipi() used atomic_clear() on _kernel.pending_ipi, which
also drained the calling CPU's own pending bit. With directed IPIs,
arch_sched_directed_ipi() skips the calling CPU, so any IPI that had
been raised for self by another CPU was silently dropped. The lost
wakeup left the target CPU stuck until something else nudged the
scheduler.
This was latent for the existing call site in ipi_work_process(), but
became reachable on every no-swap path through do_swap() once
signal_pending_ipi() was added there, producing intermittent CPU1
starvation on SMP. tests/kernel/spinlock/spinlock_api test_trylock
(and test_spinlock_bounce on cortex_a53) reproduced it reliably on
qemu_riscv32/smp, qemu_riscv64/smp, and qemu_cortex_a53/smp: CPU0's
loop would acquire bounce_lock 10000 times in a row without CPU1
ever winning a single acquisition.
Switch to atomic_and() with the self bit as the mask, which preserves
our own pending bit (so the next interrupt entry on this CPU still
sees it and runs the scheduler) while atomically draining the bits
for the other CPUs we are about to dispatch IPIs to.
Assisted-by: Claude:claude-opus-4.7
Signed-off-by: Tom Burdick <thomas.burdick@infineon.com>