The compiler cannot prove the payload copy leaves the queue fields
alone, so it reloaded the ring pointers after it. Pick the slot and
advance the pointer first.
thread_metric message processing, base / reordered / delta:
x86 133267 / 136881 / +2.71%
Xtensa DC233C 84595 / 87528 / +3.47%
RISC-V 32 80998 / 84130 / +3.87%
ARM Cortex-A53 497986 / 523815 / +5.19%
ARM Cortex-M3 96465 / 103732 / +7.53%
ARM Cortex-M4 410042 / 442637 / +7.95%
The first five are QEMU under icount; the Cortex-M4 figure is an MXChip
AZ3166 at 96 MHz. app_kernel improves 4.5% to 5.5% on one and four byte
messages and 0.5% at 192 bytes, where the copy dominates. Text shrinks
by 20 bytes and there is no impact on RAM.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
An architecture may cache the current thread in a register rather than
read it from the per-CPU structure, which is what
CONFIG_ARCH_HAS_CUSTOM_CURRENT_IMPL is for. That is worth doing on the
hottest path in the kernel, but it only pays if the compiler is allowed
to reuse a value it has already read, which in turn requires that no
code depend on _current changing underneath it.
The uniprocessor half of z_get_next_switch_handle() stores the outgoing
thread's switch handle through _current, sets the new current thread,
then reads _current again expecting the incoming one. Give it the same
treatment as its SMP counterpart and name the target once. It also stops
reading _kernel.ready_q.cache three times.
With this and the two preceding commits, nothing in the core reads the
current thread back after setting it.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
z_spin_lock_set_owner() and z_spin_unlock_valid() need the thread that
is current at the moment they run. The unlock case in particular has to
see the dummy thread that z_thread_halt() installs when an ISR aborts
the thread it interrupted, which is how it recognizes that the abandoned
lock has no owner left to match.
They read _current, which an architecture may cache in a register
(CONFIG_ARCH_HAS_CUSTOM_CURRENT_IMPL), letting the compiler carry a
value across the store that made the new thread current. Use
_raw_current, which cannot be cached. Both are called from k_spin_lock()
and k_spin_unlock() with the lock held, so interrupts are locked as that
requires.
No functional change where the architecture does not cache the current
thread, _current being _raw_current there already.
z_assert_can_swap() keeps using _current: it runs before the swap rather
than after it, and it is reachable with interrupts unlocked, where
_raw_current is not allowed.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
z_spin_lock_transfer_owner() is called from do_swap() and from
z_get_next_switch_handle() right after z_current_thread_set(), and it read
the incoming thread back from _current. An architecture that caches the
current thread in a register (CONFIG_ARCH_HAS_CUSTOM_CURRENT_IMPL) lets the
compiler reuse a value read before the store, so the lock is recorded as
still owned by the outgoing thread and the next release fails
z_spin_unlock_valid(). Reproducible under LTO.
Both call sites have the incoming thread in a local already. Pass it.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
_current is one of three things: the field the kernel records in the
per-CPU structure, a call that reads that field under an interrupt lock,
or a register an architecture keeps it in
(CONFIG_ARCH_HAS_CUSTOM_CURRENT_IMPL). Only the first is guaranteed to
report the incoming thread between z_current_thread_set() and the switch
that follows it, and code needing that has no accessor for it, so it
open-codes _current_cpu->current.
Name the field _raw_current, document what using it requires, and build
the rest on top: the uniprocessor _current is now that accessor,
z_smp_current_get() reads it under the lock it already takes, and
z_current_thread_set() writes through it.
No functional change.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
The disco_l475_iot1 is the board on which the thread_metric benchmarks
have been primarily executed.
Applying the likely() attribute to the non-SMP case where the thread
being suspended is the currently executing thread has been osbserved
to improve the thread_metric benchmark performance by about 1% on the
disco_l475_iot1 board.
This attribute resulted in the following observed changes in the
generated assembly.
1. The branch to the preferred code path is no longer a forward
branch but a fall through. CPUs typically favor fall throughs
and reverse branches over forward branches, so this helps
leverage that bias for performance.
2. The fall through to the preferred z_thread_suspend_current()
path has been observed to be in the same cache line as the
preceding instructions, whereas before this change, it was in
a different cache line. Long term, this does not guarantee that
it will always be in the same cache line as that will ultimately
depend upon the linker, but it does increase the odds.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
Return before taking usage_lock when the current CPU's accounting window
is already stopped. This avoids a redundant global-lock acquisition when
architecture IRQ-entry accounting precedes scheduler IRQ-exit accounting.
The hook already requires local interrupts to be masked, and only the
local CPU writes usage0. Keep the existing lock for every shared-counter
update and preserve runtime-statistics behavior.
Microbenchmarks on a quad-core ARMv8 SMP platform with runtime statistics
enabled showed 2.5-4.4% lower latency and up to 12.8% higher throughput.
Assisted-by: Codex:GPT-5
Signed-off-by: Aaron Wisner <aaronwisner@gmail.com>
k_pipe_write() reported K_POLL_STATE_PIPE_DATA_AVAILABLE before it put
anything into the ring buffer, and did so whether or not the write
stored any data. A thread blocked in k_poll() on
K_POLL_TYPE_PIPE_DATA_AVAILABLE was therefore woken by writes that left
the pipe with nothing to read.
The case that always misreports is an unbuffered pipe, which
k_pipe_init() accepts with a NULL buffer and a size of zero. Such a pipe
has no storage, so data only moves by direct copy to a pending reader. A
write that finds no reader transfers nothing and returns -EAGAIN, yet it
still told pollers that data was available. A write consumed entirely by
pending readers, and a write that stores nothing because the pipe is
full, misreport the same way.
Store the data first and report availability only when ring_buf_put()
accepted at least one byte. Data handed straight to a pending reader
never reaches the pipe, so it is not reported.
Enable CONFIG_POLL in the pipe test scenario. Verify that data
transferred directly from an unbuffered pipe writer to a pending reader
does not wake a poller, while data stored in a buffered pipe does.
Ran kernel.pipe.api.poll on native_sim: 21 of 21 test cases passed.
Fixes: a9ab8cb779 ("feat: enable polling support for k_pipe interface")
Assisted-by: Codex:gpt-5
Signed-off-by: Jaeyeong Lee <lee@jaeyeong.cc>
Add a common hook declaration and Kconfig switch for firmware
arguments provided by architecture startup code.
The hook receives an architecture-provided argument array and count.
This keeps the common interface independent from register naming used
by a specific architecture.
On ARM64, preserve x0-x3 from the primary CPU reset path before
startup code reuses those registers. The saved values are stored only
after the primary CPU path is selected and are passed to the hook after
RAM has been initialized for C code. Secondary CPUs do not update the
saved storage.
Signed-off-by: Vladyslav Goncharuk <vladyslav_goncharuk@epam.com>
Assisted-by: Codex:gpt-5
Add a check in thread monitor to assert if a TCB is reused while it is
still present in the kernel thread list.
This prevents undefined behavior due to TID reuse before proper cleanup.
Signed-off-by: Chidvilas Yerramsetti <cyerrams@qti.qualcomm.com>
The slim ring buffer API is now header-only and always available and the
tree has already migrated from the _claim/_finish API, so the RING_BUFFER
symbol no longer gates any code.
Remove the now-redundant "select RING_BUFFER" and CONFIG_RING_BUFFER=y
entries across drivers, subsystems, samples and tests.
Signed-off-by: Måns Ansgariusson <mansgariusson@gmail.com>
k_msgq_get() falls straight through to the empty case and leaves the
transfer behind a taken branch, though the transfer is both the common
case and the only one that copies anything.
A transfer costs five cycles less and a get on an empty queue two more,
on a path that returns without copying and that a caller hitting it
routinely would normally be blocking on rather than polling anyway.
thread_metric message processing, base / likely() / delta:
x86 133267 / 133267 / +0.00%
Xtensa DC233C 84595 / 84595 / +0.00%
ARM Cortex-M3 96465 / 97083 / +0.64%
ARM Cortex-A53 497986 / 501946 / +0.80%
RISC-V 32 80998 / 81839 / +1.04%
ARM Cortex-M4 410042 / 418985 / +2.18%
The first five are QEMU under icount; the Cortex-M4 figure is an MXChip
AZ3166 at 96 MHz, where flash wait states make a taken branch dearer.
Text shrinks by 4 bytes and there is no impact on RAM.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
Add Hexagon DSP architecture support to Zephyr RTOS, targeting the
Qualcomm Hexagon V67+ ISA running as a guest under the H2 hypervisor.
This includes:
- Architecture scaffolding (Kconfig, CMakeLists, arch selection)
- Public headers (arch.h, thread.h, exception.h, error.h)
- Core runtime: context switch, interrupt/exception handling, idle, TLS
- Memory management stubs for flat-memory H2 guest model
- SoC support for QEMU hexagon virt machine
- CMake/LLVM toolchain integration for hexagon cross-compilation
Hexagon uses H2's GEVB (Guest Event Vector Base) for interrupt dispatch
rather than a traditional function-pointer vector table, so
GEN_IRQ_VECTOR_TABLE is disabled by default.
Signed-off-by: Brian Cain <brian.cain@oss.qualcomm.com>
Previously, if a mapping operation failed partway through
k_mem_map_phys_guard(), the function would leak resources instead of
rolling back cleanly:
- map_anon_page() leaked the page frame it had just allocated when the
subsequent arch_mem_map() call failed, since it was never returned
to the free list.
- k_mem_map_phys_guard() never freed the virtual address region
obtained from virt_region_alloc() when a mapping failure occurred,
permanently leaking bitmap entries in the virtual address space.
- On partial failure, any pages already mapped before the failing one
were left mapped and, for real page frames, still marked as in use.
Now that arch_mem_map()/arch_mem_unmap() return proper error codes,
use them to roll back on failure:
- map_anon_page() frees the page frame back to the free list before
returning an error.
- A new unmap_anon_pages() helper reverses those pages previously
established by map_anon_page(): unmap and remove them from eviction
tracking if applicable, and free the page frames.
- k_mem_map_phys_guard() tracks the position of the first failed
mapping in each of its three mapping branches (demand-mapped anon,
regular anon, and known physical memory) and unwinds every
previously-succeeded page with arch_mem_unmap(), undo_anon_pages()
or undo_demand_mapped_pages() before releasing the whole virtual
region with virt_region_free().
This replaces the previous TODO referencing #28990 about rolling back
partial anonymous mappings, which is now handled.
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
Mirror the arch_mem_map() change and give arch_mem_unmap() the same
treatment: change the prototype to return int instead of void, so
architecture code has a consistent way to report unmapping failures.
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
Previously arch_mem_map() was a void function, so architecture
code had no consistent way to report mapping failures to callers
other than panicking or asserting. Change arch_mem_map() to
return int.
Note that this mostly maintains the k_panic() calls unless
the error is easy to handle (like invalid arguments). Each
architecture will need to be amended in the future to properly
handle the error conditions. Also some of the users are
converted to do k_panic() after failure. This is to keep
the signature change commit easier to review.
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Daniel Leung <daniel.leung@intel.com>
GCC cannot see that ticks_elapsed is always set before use in
z_add_timeout and fails the build with -Werror=maybe-uninitialized.
Initialize it at declaration.
Signed-off-by: Parthiban Nallathambi <parthiban@linumiz.com>
k_thread_abort() on a thread which is not the current thread takes the
immediate cleanup path in k_thread_abort_cleanup(), which calls
do_thread_cleanup(). That function ignored its thread argument and only
ever unmapped the address stashed in thread_cleanup_stack_addr, which is
populated exclusively by defer_thread_cleanup() on the self-abort path.
The mapped stack of a thread aborted by another thread was therefore
never unmapped, leaking its virtual address region plus the two guard
pages on every create/abort cycle until the virtual address space was
exhausted.
Split the two paths so each releases what it owns:
do_deferred_thread_cleanup() unmaps the stashed region, while
do_thread_cleanup() unmaps the region recorded in the thread object,
which is still valid there as the thread is neither running nor the
current thread. The latter also clears stack_info.mapped so the now
unmapped stack is no longer reachable through e.g.
k_thread_stack_space_get(), and so that a repeated cleanup of the same
thread object is a no-op.
Add a regression test which creates and aborts a thread on the same
stack object 32 times and checks that the stack is released and its
virtual address region reused on every cycle. It fails on the first
cycle without the fix.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Anas Nashif <anas.nashif@intel.com>
k_queue_append_list() notified poll waiters even when every list node
was handed directly to blocked queue consumers. The queue was empty,
but poll reported data available.
Notify poll waiters only when at least one node remains queued. Add a
regression test with two blocked consumers and one poll waiter.
The regression test fails on the parent revision and passes with this
fix.
Signed-off-by: Hui Su <3164683437@qq.com>
k_event_wait_internal() returned before acquiring the event lock when
the requested event mask was zero. As a result, reset=true did not
clear previously posted events.
Apply reset under the event lock before handling the zero-mask fast
path.
Add a regression test that verifies stale event bits are cleared.
The regression test fails on the parent revision and passes with this
fix.
Test: native_sim/native/64
- events_api: 12 passed
Signed-off-by: Hui Su <3164683437@qq.com>
k_pipe_read() and k_pipe_write() take a size_t length but report a
successful transfer as an int, and both loop until the whole request is
satisfied, accumulating across as many blocking waits as it takes. The
count is therefore not bounded by the pipe's ring buffer, and a transfer
of more than INT_MAX bytes is truncated on the way out: at exactly 2 GiB
it comes back as INT_MIN, so a caller testing for a negative result sees
a failure after every byte was moved correctly, and at 4 GiB it reports
zero.
Refuse such a request with -EOVERFLOW instead, rather than accepting it
and misreporting the outcome.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
k_sem_reset() clears the semaphore count to zero. It must not report
K_POLL_STATE_SEM_AVAILABLE to poll waiters.
The previous implementation removed and woke poll events even though no
token was available. Stop notifying semaphore poll events from
k_sem_reset(). Regular semaphore waiters still wake with -EAGAIN.
Add regression tests for single and multiple poll waiters, and for a
subsequent k_sem_give() waking a waiter after reset.
The regression tests fail on the parent revision and pass with this fix.
Test: native_sim/native/64
- poll_api: 7 passed
- poll_api_1cpu: 7 passed
Signed-off-by: Hui Su <3164683437@qq.com>
On ARMv8 a53 quad core SMP application running benchmarks, we see 54% of
sampled CPU cycles in spinlock related instructions. ~70% of these cycles
are related to the global scheduler spinlock.
Reduce global scheduler spinlock contention by short-circuiting the global
scheduler lock in the timer and IPI time-slice hook with atomic flag when
context switching is atomic and this CPU's slice has not expired.
Currently, every `z_time_slice()` acquires the scheduler spinlock, and
`z_time_slice()` is called every scheduler IPI and
`sys_clock_announce_locked()`. In practice, the vast majority of
z_time_slice calls are not time slice expiration, and `z_time_slice()` is a
NOP. Use a plain atomic flag per CPU because another CPU can publish an
expiration while the hook reads the flag without the scheduler lock.
Preserve the locked path for non-atomic context switches, including
pending_current synchronization. Keep expired-slice processing, per-thread
callbacks, round-robin rotation and timeout reset behavior.
We observe a 2.9% to 4.7% latency improvement and up to 10.8% throughput
improvement.
Assisted-by: Codex:GPT-5
Signed-off-by: Aaron Wisner <aaronwisner@gmail.com>
A thread actively running on another CPU at panic time was previously
invisible to real debugging: coredump only captured the panicking CPU's
exception frame, so every other thread's registers came from its
k_thread.callee_saved struct -- correct for genuinely sleeping threads,
but stale (reflecting whenever it last voluntarily context switched) for
one that's actually running elsewhere right now.
Add CONFIG_DEBUG_COREDUMP_SMP_FREEZE_CPUS (default y where the arch sets
ARCH_SUPPORTS_COREDUMP_SMP_FREEZE, i.e. arm64 Cortex-A SMP). On panic,
coredump() freezes every other online CPU before dumping and thaws them
after the memory-region walk, and emits each frozen CPU's live registers
as a new coredump section (COREDUMP_CPU_SNAPSHOT_HDR_ID, tagged with CPU
index and thread pointer, reusing the arch-info block register layout).
The panicking CPU (never frozen, since freeze only targets other CPUs)
emits a zero-payload marker recording only which thread was panicking
there, which the host side uses to identify it unambiguously.
arm64 implementation (arch/arm64/core/smp.c, coredump.c): send a new SGI
to every other online CPU. Each handler captures its exact live state --
x0-x18/lr/spsr/elr via the arch_esf that _isr_wrapper() stashes at a
fixed, nesting-independent slot on its IRQ stack, plus x19-x29 via inline
asm (valid per AAPCS64) -- then spins holding whatever it held before
being frozen, untouched, until the dump finishes, and resumes via the
standard exception-return path exactly where it was interrupted. This
matters because a panic here can be recoverable (this tree's own
coredump_threads test doesn't halt/reboot after dumping).
The SGI uses a deliberately low priority (SGI_COREDUMP_FREEZE_PRIO), not
IRQ_DEFAULT_PRIORITY: a frozen CPU never returns from its handler so
_isr_wrapper() never EOIs, and per GIC priority rules only a strictly
higher-priority interrupt can preempt a still-active one. At the default
priority a frozen CPU could not service same-priority peripheral
interrupts (confirmed on hardware as Ethernet TX-done starvation during
the UDP backend); the lowest usable priority lets everything else keep
making progress while a CPU waits.
The freeze/thaw handshake uses two separate per-CPU signals: freeze_state
(IDLE/REQUESTED/CAPTURED) and a dedicated release_requested flag that
only thaw writes. Kept separate so a late-arriving CPU (whose SGI was
delayed past the freeze-side timeout) still sees a release signal instead
of stomping it with its own CAPTURED update and spinning forever; thaw
sets release_requested unconditionally for every other CPU. The handler
wait is also bounded by a 60-second wall-clock backstop (k_cycle_get_32(),
safe from any context) so no CPU is ever stuck indefinitely. A CPU that
never responds within the freeze-side timeout (never booted, or stuck
with IRQs masked) is skipped and never blocks the dump.
The host coredump log parser learns to parse the new snapshot section.
Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
This removes k_futex from the kobject table and allows any address
accessible to userspace to be used as futex and fixes#25497.
Instead of different wait queues per futex, this solution uses a
single wait queue for all threads waiting on any futex. The futex
address is stored in the thread structure. On wake, the kernel
walks through the queue and wakes either the highest priority or
all threads waiting on the specified futex address.
This is more lightweight than #42087 and more efficient for smaller
numbers of threads, but does not scale well to large numbers of
threads waiting on different futexes.
Add a test for the correct order when waking only one thread.
Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
Add CONFIG_IPI_OPTIMIZE_IDLE to avoid waking multiple idle CPUs for a
single runnable thread. Select one active, affinity-compatible idle CPU
and record which thread its outstanding scheduling opportunity covers.
Track coverage with a reserved CPU bitmap and a per-CPU target array.
Each idle CPU and runnable thread can have at most one reservation. The
mapping represents coverage rather than binding: any CPU may select any
eligible runnable thread.
When a reserved CPU selects a different thread, transfer the selected
thread's existing reservation to the original target when possible.
Otherwise, issue a replacement IPI for the still-runnable target. Clear
reservations when their target leaves the run queue.
Apply the same optimization to MetaIRQ wakeups. The MetaIRQ can preempt
its local waker while one idle CPU is notified to run the displaced
preemptible work.
Signed-off-by: Jinming Zhao <jinmzhao@qti.qualcomm.com>
sys_sflist stores flags in unused low bits of node links. The number of
available bits follows the ABI's natural pointer alignment, but the generic
header currently forces pointer-size alignment and globally requires two
bits.
That guarantee does not hold for borrowed nodes such as intrusive k_queue
items, which are caller-owned objects cast to sys_sfnode_t. On ABIs with
four-byte pointers and two-byte pointer alignment, only bit 0 is reliably
free.
Remove the forced alignment, expose SYS_SFLIST_FLAG_BITS, and let consumers
assert their own requirements. k_queue and the MMU free-page list each
require one bit. Make the unit test honor the capacity reported by the ABI
and clarify the intrusive queue alignment documentation.
This preserves generic flagged lists and safely supports one-bit ABIs.
Assisted-by: ChatGPT:gpt-5.6
Signed-off-by: Dimitri Varpusvuori <dimitri.varpusvuori@gmail.com>
This commit introduces a new timeout backend using a
Pugh skip list, which allows for O(log N) insertion
and removal of timeouts. The skip list maintains
a same-tick firing order and has no capacity
limit, making it suitable for systems with many
concurrent timeouts. The implementation includes
necessary Kconfig options and updates to the timeout
management code to support this new backend.
Signed-off-by: Dhruv Menon <dhruvmenon1104@gmail.com>
Replace the system-wide semaphore spinlock with a configurable array
of naturally aligned lock stripes. Hash each semaphore by its natural
alignment without changing the semaphore object ABI.
Default CONFIG_SEM_LOCK_STRIPES to one so SMP and uniprocessor builds
preserve existing global-lock behavior and static memory use. Use a
scalar lock unless SMP striping is explicitly enabled, so UP builds
retain zero-sized spinlocks and avoid non-portable empty-struct arrays.
Require a power-of-two stripe count so the compiler reduces stripe
selection to a bitmask. Divide by the alignment of struct k_sem before
hashing to discard redundant address bits.
Do not impose cache-line alignment or padding. Explain how spinlock
and cache-line sizes affect the stripe count needed to reduce cache
contention.
Validate upstream semaphore regressions with zero-sized uniprocessor
spinlocks and SMP configurations using one, 16, and 32 stripes.
Reject non-power-of-two stripe counts at compile time.
An internal throughput microbenchmark on a quad-core ARM Cortex-A53
SMP system observed an improvement of approximately 3.1%.
Assisted-by: Codex:GPT-5
Signed-off-by: Aaron Wisner <awiz@openai.com>
The SoC and board hooks are selected by the
SoC or the board. They also implemt them.
Remove the ability to enable them regularly
by the user as the build would then fail, because
the board or SoC has not implemented them.
Signed-off-by: Fin Maaß <f.maass@vogl-electronic.com>
Remove remaining uses of the internal __ASSERT_ON macro. Let __ASSERT()
handle disabled assertions, mark assert-only values as unused where needed,
and use CONFIG_ASSERT for assertion-only state.
Signed-off-by: Måns Ansgariusson <mansgariusson@gmail.com>
Remove reference to legacy api, where the delay
is fixed to K_FOREVER, as that, since quite some time,
no longer exists.
Signed-off-by: Fin Maaß <f.maass@vogl-electronic.com>
Replace the unsigned-arithmetic range checks in k_mem_map_phys_bare()
and k_mem_unmap_phys_bare() with the equivalent overflow-free form:
__ASSERT(aligned_size - 1 <= (UINTPTR_MAX - aligned_addr), ...)
The original checks were correct, but the explicit form avoids
relying on unsigned wraparound behavior and is consistent across
all assert sites in mmu.c.
Signed-off-by: Lokesh Gundu <lgundu@qti.qualcomm.com>
An empty k_queue_get() emits the blocking trace before checking
K_NO_WAIT. This records a block that never occurs and invokes the
tracing backend on a non-blocking fast path.
Move the trace below the no-wait return so only calls that can pend
emit it. Add CTF coverage for an empty no-wait get.
Signed-off-by: Magnus Strømme <magnus.henrik@hotmail.com>
k_stack_alloc_init() calculates the allocation size as
num_entries * sizeof(stack_data_t). On targets where size_t is 32 bits,
a sufficiently large entry count can overflow the multiplication and
result in an allocation much smaller than the capacity recorded in the
stack object.
The userspace verifier already checks this multiplication, but supervisor
callers invoke z_impl_k_stack_alloc_init() directly and bypass that check.
Check the multiplication in the implementation before allocating the
buffer. Treat an overflowing request as an allocation failure and return
-ENOMEM.
Add regression coverage for the supervisor allocation path. The complete
qemu_cortex_m3 stack validation passed after the fix:
16 tests passed, 0 failed
PROJECT EXECUTION SUCCESSFUL
Fixes: be3d4232c2 ("kernel: fix k_stack_alloc_init()")
Signed-off-by: Hui Su <3164683437@qq.com>
k_thread_stack_alloc() passed K_KERNEL_STACK_LEN(size) directly to the
allocator. The alignment and reserved-size arithmetic could wrap for an
oversized request, causing it to be turned into a much smaller allocation
instead of being rejected.
Before the fix, the regression test accepted SIZE_MAX and reported:
Assertion failed ... (stack is not NULL)
overflowing stack size must be rejected
FAIL - test_dynamic_thread_stack_size_overflow
Reject the allocation when K_KERNEL_STACK_LEN(size) is smaller than the
requested size, which indicates that the size calculation wrapped. The
existing userspace dynamic-object overflow check remains unchanged.
Fixes: 7b1b2576ac ("kernel: support dynamic thread stack allocation")
Signed-off-by: Hui Su <3164683437@qq.com>
k_queue_alloc_append() and k_queue_alloc_prepend() wrap the caller's data
in an internal alloc_node before inserting it into the queue. The old
k_queue_remove() implementation searched for the caller data pointer as
if it were the linked-list node, so removal returned false and left the
allocated queue item in place.
Walk the queue and compare the unwrapped payload pointer instead. Free
the alloc_node wrapper after removing a matching item, while preserving
the existing behavior for directly inserted queue nodes.
Add coverage for removing items inserted through both alloc APIs.
Tested:
- ZEPHYR_TOOLCHAIN_VARIANT=gnuarmemb GNUARMEMB_TOOLCHAIN_PATH=/usr
west build -d /tmp/zephyr-queue-gnuarmemb
- qemu_cortex_m3: queue_api 15/15 and queue_api_1cpu 7/7
Signed-off-by: Hui Su <3164683437@qq.com>
Splits z_waitq_head() into two versions: z_waitq_head_locked() and
z_waitq_head(). When the scheduler's spinlock is known to be already
held, z_waitq_head_locked() should be used--otherwise, z_waitq_head()
is to be used.
However, this approach uncovered a path where the scheduler spinlock
could be recursively taken when a thread is aborted. To work around the
recursion (see k_thread_perms_all_clear), knowledge of the scheduler's
spinlock state must be passed to the lower layers for use in the
cleanup routines for message queues, stacks and timers.
Fixes#115756
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
As k_malloc(), k_calloc(), k_realloc() are non-blocking calls,
k_free() should not be calling z_waitq_head() to look for blocked
threads. Developers that want the blocking/waking capabilities
should use the k_heap_xxx() routines instead.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
Updates the k_stack_cleanup() and k_msgq_cleanup() routines
to lock the appropriate lock to prevent two execution contexts
from trying to clean up the same object at the same time.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
As part of an effort to abstract away the use of _sched_spinlock in
the kernel, this commit introduces z_reschedule_locked(). Callers
must already have _sched_spinlock held.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
As part of an effort to abstract away the use of _sched_spinlock in
the kernel, this commit introduces z_swap_locked(). Callers must
already have _sched_spinlock held.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The scheduler's spinlock was originally confined to sched.c. However,
as the kernel has evolved and functionality has been moved around,
not only have its references proliferated, but the header files that
reference it have become somewhat brittle. This commits adds a new
private header that will abstract away most of the scheduler's
spinlock references to help keep things cleaner and applies them to
the kernel.
Signed-off-by: Peter Mitsis <peter.mitsis@intel.com>
The sleep trace hooks were attached to k_sleep(), k_msleep() and
k_usleep(). Those are inline wrappers now, and the only function left to
instrument is the k_sleep_ticks() primitive underneath them, so the three
hook families no longer have anything emitting them. The msleep pair had
in fact been unused for some time already.
Replace all three with one k_thread_sleep_ticks pair reporting the timeout
and the time left to sleep in ticks, and drop what is left behind in each
backend: the CTF events and their top level helpers, the entry in
SYSVIEW_Zephyr.txt, and the user and test hooks. The new CTF and
SystemView ids are fresh rather than reused, so a recording made by an
older build cannot be misread by a newer decoder. The retired SystemView
ids stay defined, since out of tree code may refer to them.
The profiling example in the instrumentation documentation named two
symbols that no longer exist, and is updated to the ones that replace
them.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
k_sleep() and k_usleep() resolve to an out of line implementation living in
another translation unit, a system call under CONFIG_USERSPACE and a plain
call otherwise. Either way the compiler sees neither whether the requested
duration is a constant nor whether the caller cares about the returned
value, so it always emits the tick to millisecond and microsecond
conversions. Those conversions are the only thing that drags the 64 bit
division helper into many small builds: CONFIG_SYS_CLOCK_TICKS_PER_SEC
defaults to 10000 on tickless platforms, which lands in the integer
division path and divides a 64 bit value by ten, something GCC will not
strength reduce on a 32 bit target.
Promote the former z_tick_sleep() helper to a k_sleep_ticks() system call
that reports the remaining time in ticks, and rebuild k_sleep() and
k_usleep() as inlines on top of it. The conversions now live at the call
site where the compiler can fold or discard them. Nearly every caller
discards them: of the 4036 sleep call sites in the tree, exactly six use
the returned value, and only three of those are outside of tests.
What this removes is best seen in the two out of line entry points that
disappear, which on a nucleo_f030r8 (Cortex-M0, 10000 ticks/s) were the
only two callers of the 64 bit division helper in the whole image:
08002594 <z_impl_k_sleep>:
bl 8001... <z_tick_sleep>
...
movs r0, #9 /* the ceil() bias, 10 - 1 */
adds r0, r0, r2
adcs r1, r3
movs r2, #10 /* ticks / 10 -> milliseconds */
bl 8000180 <__aeabi_uldivmod>
080025c0 <z_impl_k_usleep>:
movs r0, #99 /* the ceil() bias, 100 - 1 */
...
movs r2, #100 /* microseconds / 100 -> ticks */
bl 8000180 <__aeabi_uldivmod>
bl 8001... <z_tick_sleep>
Both are gone, and so are __aeabi_uldivmod and __udivmoddi4, which are no
longer referenced anywhere. A k_usleep(250) call site now passes the
literal 3 ticks the compiler worked out on its own, where it used to pass
250 microseconds to a helper that divided by 100 at runtime. For a loop
calling k_msleep(100), k_sleep(K_SECONDS(1)) and k_usleep(250) without
using the results, FLASH drops from 11028 to 10600 bytes.
No trace point is emitted here. The existing sleep hooks take a duration
in milliseconds or microseconds, which no longer describes what this
function is handed, and tracing cannot live in the inlines either: every
tracing backend header includes kernel.h, so a translation unit reaching a
backend header first would expand SYS_PORT_TRACING_* before those macros
exist. The following commit gives k_sleep_ticks() a hook of its own.
The duration goes through Z_TIMEOUT_US() rather than an open coded
conversion. That clamps a negative argument to zero, which k_usleep()
never did: it used to truncate the 64 bit conversion into an int32_t
first, which hid the worst of it, and passing it through whole would turn a
negative duration into a sleep of roughly 1.8e17 ticks rather than a bogus
absolute deadline. The macro also carries the cast to k_ticks_t that C++
needs, since Z_TIMEOUT_TICKS_INIT() is a braced initializer there and an
unsigned to signed conversion inside one is a narrowing error.
This also removes a potential link error. k_usleep() had no
implementation at all with CONFIG_MULTITHREADING=n, since nothread.c only
ever provided z_impl_k_sleep(), so any caller failed to link. Nothing in
the tree happens to call it in that configuration, which is presumably why
it went unnoticed; a single z_impl_k_sleep_ticks() now serves both.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
k_sleep(), k_msleep() and k_usleep() are declared in kernel.h today.
Later commits grow that API with a tick based primitive plus inline unit
conversions, which is more than kernel.h ought to carry for a single
service.
Move the three declarations verbatim into a new include/zephyr/sleep.h
and have kernel.h include it, so existing users need no change. The new
header has to be registered with zephyr_syscall_header(), otherwise the
marshalling stubs for k_sleep() and k_usleep() stop being generated and
CONFIG_USERSPACE builds fail to find k_sleep_mrsh.c.
The unit_testing platform fabricates empty stand ins for the generated
syscall headers rather than running the generator, so sleep.h has to be
added to that list too. Without it any unit test reaching kernel.h fails
to find zephyr/syscalls/sleep.h.
The sleep syscalls also have to be excluded from the automatic syscall
tracing in gen_syscalls.py, as kernel.h already is. They carry hand
written trace points in kernel/sleep.c, and without the exclusion the
generated wrapper would wrap every call in a second set.
No functional change. A test application built for nucleo_f030r8 comes
out byte identical, at 10952 bytes of FLASH before and after.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
sys_rand_get() is used to initialize __stack_chk_guard under
CONFIG_STACK_CANARIES_TLS, but <zephyr/random/random.h> was included
under CONFIG_CURRENT_THREAD_USE_TLS. Builds fail when
STACK_CANARIES_TLS=y and CURRENT_THREAD_USE_TLS=n. Also fix the
mismatched endif comment.
Signed-off-by: Yiren Guo <guoyr_2013@hotmail.com>
With CONFIG_SCHED_THREAD_USAGE_ANALYSIS=y and
CONFIG_SCHED_THREAD_USAGE_ALL=n, sched_cpu_update_usage() is a no-op
macro, so the cycles local in k_thread_runtime_stats_enable() and
z_thread_stats_reset() is set but never read, triggering a compiler
warning. Pass the expression straight to the macro to keep
that configuration warning-free.
Signed-off-by: Yiren Guo <guoyr_2013@hotmail.com>