Commit graph zephyr/arch
Author SHA1 Message Date
Joel Holdsworth
f06b56caa7 arch: hexagon: support building without generated SW ISR table
Guard references to _sw_isr_table in irq_manage.c so that applications can
build without generated ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Joel Holdsworth
e3802d3e1f arch: riscv: support building without generated SW ISR table
Make GEN_SW_ISR_TABLE depend on GEN_ISR_TABLES, and guard the import of
_sw_isr_table in isr.S so that applications can build without generated
ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Joel Holdsworth
a7bc5ebbc9 arch: rx: support building without generated SW ISR table
Make GEN_SW_ISR_TABLE depend on GEN_ISR_TABLES, and guard references to
_sw_isr_table in vects.c so that applications can build without generated
ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Joel Holdsworth
cf3fe97cd8 arch: mips: support building without generated SW ISR table
Make GEN_SW_ISR_TABLE depend on GEN_ISR_TABLES, and guard references to
_sw_isr_table in irq_manage.c so that applications can build without
generated ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Joel Holdsworth
8173e314bb arch: sparc: support building without generated SW ISR table
Make GEN_SW_ISR_TABLE depend on GEN_ISR_TABLES, and guard references to
_sw_isr_table in irq_manage.c so that applications can build without
generated ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Joel Holdsworth
29ffa7e58d arch: openrisc: support building without generated SW ISR table
Make GEN_SW_ISR_TABLE depend on GEN_ISR_TABLES, and guard references to
_sw_isr_table in irq_manage.c so that applications can build without
generated ISR tables when interrupts are unused.

Assisted-by: Gemini:gemini-3.8-flash
Signed-off-by: Joel Holdsworth <jholdsworth@nvidia.com>
2026-09-26 08:40:33 +02:00
Sefa Celik
e560c91b1f arch: riscv: andes: handle NMI arriving at the reset vector
On AndeStar V5 cores such as the N25/N25F, the NMI vector base address
register (mnvec, CSR 0x7C3) is read-only and reads back the reset_vector
input signal. An NMI therefore enters the system at __reset, and the
boot path re-initialises the core, destroying the machine state the NMI
was raised to let software examine.

Add CONFIG_RISCV_CUSTOM_CSR_ANDES_NMI, gated on a CPU_HAS_ANDES_NMI
capability that the SoC selects, so the option is only offered on cores
implementing the Andes NMI extension rather than the RISC-V RNMI
extension. It tells the two cases apart at the top of __reset and
branches to _andes_nmi_entry instead of booting. Coming out of reset
mcause and mepc both read 0, whereas an NMI sets mcause to 1 and leaves
the interrupted PC in mepc. The option defaults to n, so behaviour is
unchanged unless a SoC enables it.

Note that Andes reports the NMI with the mcause interrupt bit clear,
rather than set as the privileged specification recommends, so the check
compares mcause against 1 exactly. That value is also the instruction
access fault code. Such exceptions are taken through mtvec, which is set
up early in boot, so in practice they do not reach the reset vector;
mtvec.BASE does reset to 0 however, so a fault raised before that setup
cannot be told apart from an NMI.

_andes_nmi_entry is weak and only parks the core, so that the option
links on its own. A SoC enabling NMIs is expected to provide a real
handler overriding it, and to set up whatever that handler needs, such
as an NMI stack in mscratch, before enabling any NMI source.

Signed-off-by: Sefa Celik <sefa.celik@analog.com>
2026-09-26 01:14:25 +02:00
Aaron Wisner
14ee3d6411 arch: arm64: cache the current thread in TPIDR_EL1 on Cortex-A SMP
Use the existing custom-current architecture hooks to keep the running
thread pointer in TPIDR_EL1. A register read avoids the call and IRQ
masking needed to prevent migration between a CPU lookup and the
current-thread load.

Keep the per-CPU current pointer for cross-CPU readers. The scheduler
updates the register with preemption disabled, and reset initializes it
before the kernel installs the dummy thread on each CPU. TPIDR_EL0
remains available for TLS and TPIDRRO_EL0 for CPU and exception state.

Enable this for Cortex-A SMP without system power management.
Uniprocessor and PM builds keep the existing lookup; power-state
restoration is not extended to save the register.

A53 quad core microbenchmarks show ~1.2% higher throughput and up
to 4.7% lower latency.

Assisted-by: Codex:GPT-5
Signed-off-by: Aaron Wisner <aaronwisner@gmail.com>
2026-09-26 00:48:21 +02:00
Benjamin Cabé
8b93109f26 arch: arm: cortex_m: pend PendSV with a plain store on thread abort
k_thread_abort() pends PendSV with SCB->ICSR |= PENDSVSET, the third
such site after arch_swap() and z_arm_exc_exit(). ICSR's writable bits
are write-one-to-set or write-one-to-clear and writing zero is a no-op,
so the read-modify-write is equivalent to a plain store while also
writing back a stale snapshot of the other write-one bits.

This path is not hot, so the motivation is consistency and dropping
that stale write-back rather than the saved access. The SHCSR update
immediately below stays a read-modify-write: its bits are ordinary
read/write ones that have to be preserved.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
2026-09-26 00:44:46 +02:00
Benjamin Cabé
6f0188a56c arch: arm: cortex_m: pend PendSV with a plain ICSR store
arch_swap() and z_arm_exc_exit() pended PendSV with a read-modify-write
of SCB->ICSR. The register's writable bits are all write-one-to-set or
write-one-to-clear and writing zero to them has no effect, so the read
and the OR are pointless: a plain store of PENDSVSET is equivalent and
one memory access shorter on the two hottest exception paths. It also
avoids writing back a stale snapshot of the other write-one bits (for
instance re-pending a SysTick whose pend bit was cleared between the
read and the write, were the sequence ever preempted).

Measured standalone against main (latency_measure in cycles, lower is
better; thread_metric scores, higher is better):

* mps2/an385 (Cortex-M3, QEMU icount):
  - latency_measure: mean -1.8% over 47 ops, min/max/median
    -5.0/+0.0/-1.1%; k_yield context switch 182 -> 180 cycles.
  - thread_metric: cooperative and preemptive +1.1%.

* az3166_iotdevkit (STM32F412, Cortex-M4 @ 96 MHz):
  - thread_metric: preemptive +2.0%, cooperative +1.1%.
  - latency_measure: ops stay within ±2 cycles of code-placement
    noise, min/max/median -2.0/+2.2/+0.0%.

* Flash, az3166 latency_measure image: -8 B at both -Os and -O2.

Assisted-by: Claude:fable-5
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
2026-09-26 00:44:46 +02:00
Chidvilas Yerramsetti
014488bbc2 cmake: arm: clang: add -mtp=soft for thread-local storage
Clang's default ARM TLS codegen accesses the TLS block through
TPIDRURO (the per-CPU pointer) instead of TPIDRURW, which Zephyr
uses as the TLS base pointer on Cortex-A/R. This produced corrupted
z_tls_current and other thread-local accesses since the wrong base
was used.

Pass -mtp=soft to force calls through Zephyr's own __aeabi_read_tp,
matching the GCC toolchain's target_arm.cmake and Zephyr's TPIDRURW
based TLS mechanism.

Signed-off-by: Chidvilas Yerramsetti <cyerrams@qti.qualcomm.com>
2026-09-26 00:44:34 +02:00
Benjamin Cabé
143d9e91fd arch: tricore: lock interrupts across self-abort
An interrupt taken between setting to_reclaim and z_thread_abort()
marking the thread dead made isr_wrapper reclaim the CSAs of a thread
that was still runnable. Switching back to it trapped with a call stack
underflow (seen in benchmark.posix.threads on qemu_tc3x).

Lock interrupts before setting to_reclaim.

Assisted-by: Claude:opus-5.5
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
2026-09-25 11:33:25 -05:00
Fin Maaß
80c6a6f2ba arch: riscv: grant S-mode access to the Zkr seed CSR
The seed CSR of the Zkr extension is only accessible from M-mode after
reset. When the kernel runs in S-mode on top of the in-tree M-mode
runtime, set mseccfg.SSEED before dropping to S-mode, so that the
kernel can use the entropy source. U-mode access stays disabled.

With an external SBI implementation this is up to the firmware, OpenSBI
sets SSEED on harts that implement Zkr.

Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Fin Maaß <f.maass@vogl-electronic.com>
2026-09-24 17:12:58 +02:00
Fin Maaß
d4120a7b0d arch: riscv: add Zkr ISA extension support
Add RISCV_ISA_EXT_ZKR for the Zkr entropy source extension. As for the
other ISA extensions, it is enabled from the riscv,isa-extensions
devicetree property. It is also enabled by Zk, which includes Zkr.

Add the addresses of the seed and mseccfg CSRs and the mseccfg.SSEED
bit, which grants S-mode access to the seed CSR.

Zkr is not added to -march: it adds no instructions and the CSRs are
addressed by number, while the LLVM toolchain of the Zephyr SDK selects
its multilib from the exact -march string and has none matching zkr.

Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Fin Maaß <f.maass@vogl-electronic.com>
2026-09-24 17:12:58 +02:00
Sylvio Alves
699c613f07 arch: xtensa: carry backtrace mask and cause in the frame
The window-increment mask and the exception cause were kept in
file-scope statics written only by xtensa_backtrace_print(). Two
cores panicking at once clobbered each other's values, and a
direct caller of the public xtensa_backtrace_get_next_frame() ran
with a stale mask and got wrong PCs.

Move both into struct xtensa_backtrace_frame_t and pass them to
the PC fixup, making it a pure function. Output is unchanged.

Assisted-by: Claude:opus-5
Signed-off-by: Sylvio Alves <sylvio.alves@espressif.com>
2026-09-24 09:57:04 +02:00
Ayush Singh
c847bfbc8f arch: riscv: Add initial smp support with external sbi
This patch only allows enabling SMP in 1 core setups.

Signed-off-by: Ayush Singh <ayush@beagleboard.org>
2026-09-24 03:49:59 +02:00
Ayush Singh
27619328e7 arch: riscv: core: Add ipi_sbi support
RISC-V Supervisor Binary Interface Specification supports
inter-processor interrupts using the IPI extension.

Signed-off-by: Ayush Singh <ayush@beagleboard.org>
2026-09-24 03:49:59 +02:00
Alberto Escolar Piedras
d123b55e92 arch/posix: Export native target architecture as cache variable
Save the target architecture as a cmake cache variable, so other
components can adapt their build if they depend on it.

Note this target architecture is not necessarily the host processor
as posix arch based targets can also be crosscompiled.

Signed-off-by: Alberto Escolar Piedras <alberto.escolar.piedras@nordicsemi.no>
2026-09-23 19:20:33 +02:00
Abderrahmane JARMOUNI
06520865aa arch: Kconfig: improve XIP help text
Clarify the CONFIG_XIP description to explain its effect on LMA/VMA
alignment for .text/.rodata and RAM usage, and note that it should be
disabled for images copied into RAM before execution.

Signed-off-by: Abderrahmane JARMOUNI <git@jarmouni.me>
2026-09-21 16:28:51 -04:00
Jamie McCrae
1909f76910 kconfig: Correctly have EXPERIMENTAL prompt
Fixes various issues with Kconfigs, including:
  - Using the wrong indentation
  - Not having [EXPERIMENTAL] in the prompt, or not having it at
    the end
  - Wrongly stating the CONFIG_EXPERIMENTAL is needed to enable
    an experimental Kconfig

Signed-off-by: Jamie McCrae <jamie.mccrae@nordicsemi.no>
2026-09-21 16:27:52 -04:00
Hongquan Li
3bedc2935a arch: arc: mpu: fix region index in remove_mem_partition
arc_core_mpu_remove_mem_partition() disabled the region at
base_region + partition_id, counting upward from the domain-partition
base region, while arc_core_mpu_configure_mem_domain() assigns array
slots counting downward (partitions[i] -> region base - i) and, since
the hole-skip fix, skips zero-sized entries. The two disagree even
without holes, so the wrong region would be disabled if the function
were ever wired up.

Derive the region slot from the partition's array position: count the
valid partitions before partition_id and subtract from the base region,
matching how the configure loop programs them. No-op on holes and
out-of-range ids.

The function has no in-tree callers today; this aligns it with the
configure-side mapping so it is correct if it is ever used.

Fixes: #116321

Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
2026-09-19 10:30:13 +02:00
Hongquan Li
a39ff708e3 arch: arc: mpu: skip zero-sized holes in memory domain partitions
k_mem_domain_remove_partition() zeroes a partition slot in place and
decrements num_partitions, leaving a hole in partitions[]. The ARC MPU
domain region programming loops bounded their iteration with
num_partitions, so each hole consumed one iteration and any valid
partition behind it was never programmed into the MPU, faulting
user-mode accesses with EV_ProtV.

Scan the whole partitions array (bounded by
CONFIG_MAX_DOMAIN_PARTITIONS, since num_partitions no longer reflects
the number of slots to inspect) and program only the non-zero-sized
entries, decrementing the remaining-partition counter only for valid
ones. This matches the arm32 implementation in arm_core_mpu.c and is
applied in:

- arc_core_mpu_configure_mem_domain() (MPU v2/v3/v6, common header)
- arc_core_mpu_configure_thread() gap-filling path (MPU v4/v8)
- arc_core_mpu_configure_mem_domain() non-gap-filling path (MPU v4/v8)
- arc_core_mpu_remove_mem_domain() (MPU v4/v8)

Fixes: #116321

Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
2026-09-19 10:30:13 +02:00
Vladyslav Goncharuk
43774bbb43 arch: add firmware argument handler hook
Add a common hook declaration and Kconfig switch for firmware
arguments provided by architecture startup code.

The hook receives an architecture-provided argument array and count.
This keeps the common interface independent from register naming used
by a specific architecture.

On ARM64, preserve x0-x3 from the primary CPU reset path before
startup code reuses those registers. The saved values are stored only
after the primary CPU path is selected and are passed to the hook after
RAM has been initialized for C code. Secondary CPUs do not update the
saved storage.

Signed-off-by: Vladyslav Goncharuk <vladyslav_goncharuk@epam.com>
Assisted-by: Codex:gpt-5
2026-09-18 11:14:23 +01:00
Christoph Seitz
bc878d271b arch: add TriCore architecture support
Add TriCore architecture port supporting TC1.6P and TC1.8P core
variants used in Infineon AURIX TC3x and TC4x SoC families.

Includes context switching using TriCore upper/lower context save
areas (CSA), interrupt management, trap and exception handling,
syscall interface, reset vector and linker script.

Supports cooperative and preemptive threading, nested interrupt
handling with priority-based preemption.

Signed-off-by: Christoph Seitz <christoph.seitz@infineon.com>
Signed-off-by: Parthiban Nallathambi <parthiban@linumiz.com>
2026-09-18 11:13:16 +01:00
Abderrahmane JARMOUNI
a399a056ab arch: arc: core: rework usage of cache line size API
A component should not rely on the API it is implementing.
In this case, the arch layer is implementing the arch cache API
(include/zephyr/arch/cache.h), that is used by the public sys cache
API (include/zephyr/cache.h), so it can't call the latter.

Also rework init_dcache_line_size logic since now
CONFIG_DCACHE_LINE_SIZE is always defined even if DCache line size
runtime detection is available.

Signed-off-by: Abderrahmane JARMOUNI <git@jarmouni.me>
2026-09-17 12:19:21 +01:00
Shubhankar Kulkarni
c1c9d40b8b arch: arm64: prevent unaligned SIMD accesses in pre-MMU code with clang
Before z_arm64_mm_init() enables the MMU, all accesses are treated as
Device-nGnRnE, which the architecture requires to be naturally aligned.
This is independent of SCTLR_ELx.A: clearing the A bit relaxes alignment
checking for Normal memory only, it does not make unaligned Device
accesses legal.

Clang may lower struct assignments and small memory copies in this
library into 128-bit Advanced SIMD accesses (ldp/stp q, ldur/stur q).
When such an access is not 16-byte aligned and executes before the MMU
is enabled, it raises an alignment fault (data abort, ESR_ELx.DFSC
0b100001). As this happens before console initialisation, the failure
presents as a silent hang.

GCC already avoids this instruction class here via
-moverride=tune=no_ldp_stp_qregs, but there is no clang equivalent, and
clang rejects that argument. That option would also be insufficient on
its own, as it only suppresses ldp/stp pairs and not a single unaligned
128-bit access.

Apply -mstrict-align to this library when building with clang. The flag
is scoped with zephyr_library_compile_options() rather than
zephyr_cc_option() so that code running with the MMU enabled is not
restricted. Assembly is unaffected, so the FPU context save/restore in
fpu.S continues to use the Qn registers.

Signed-off-by: Shubhankar Kulkarni <shukul@qti.qualcomm.com>
2026-09-17 12:12:59 +01:00
Laurie Fay
d7b131763f arch: arm: Use CMSIS APIs for USE_SWITCH
Replace hand-written assembly with CMSIS APIs in the implementation
of USE_SWITCH, as in the review comments for #85248. This should
allow better maintainability and compiler compatibility, and ensure
correctness with respect to instruction barriers.

A small change to include order is needed to avoid errors when the
cmsis_core.h header includes zephyr/irq.h before the
arch/arm/irq.h implementation is included.

Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Laurie Fay <Laurie.Fay@arm.com>
2026-09-17 12:10:55 +01:00
Silesh C V
afaff5491e arch: arm: aarch32: enable SMP for ARMv8-A AArch32
The shared Cortex-A/R SMP support (GIC SGI-based scheduler IPIs,
secondary core bring-up via reset-time voting locks) is sufficient to
run SMP on ARMv8-A AArch32. Enabling it only requires selecting the
capability flags SCHED_IPI_SUPPORTED (under SMP) and
ARCH_HAS_DIRECTED_IPIS. Also drop the stale "UP only" note.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
046c597f13 arch: arm: cortex_a_r: smp: skip unmapped CPUs when sending IPIs
send_ipi() means to skip CPUs whose cpu_map[] entry is not yet
populated, but it compares the sender's MPIDR against INV_MPID
instead of the target's. GET_MPIDR() never returns INV_MPID, so the
guard never fires and an unmapped entry reaches gic_raise_sgi().
Fix this.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
79c794c6f8 arch: arm: remove unused offload_routine extern declarations
Remove unused extern declarations of offload_routine from cortex_m
and cortex_a_r exception.h as currently there are no users for these.

The global offload_routine (and the irq_offload code that set it to
the offloaded routine and cleared it back to NULL) was introduced by
commit 75caa2b084 ("arm: exception-assisted kernel panic/oops support")
so that _IsInIsr() could tell whether the active SVC was an irq_offload()
call.

On cortex-m, commit 4f11b6f8cf ("arch: arm: re-implement
z_arch_is_in_isr") switched arch_is_in_isr() to read the IPSR and
dropped the offload_routine != NULL check.

On cortex_a_r, the declaration was introduced by commit c30a71df95
("arch: arm: Add Cortex-R support") but it was never referenced.
arch_is_in_isr() derives interrupt context from arch_curr_cpu()->nested.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
1bca4d22dc arch: arm: cortex_a_r: add SMP/nested-capable irq_offload
The Cortex-A/R cores previously shared the Cortex-M irq_offload
implementation (arch/arm/core/irq_offload.c), which stores the offloaded
routine and parameter in global state and brackets the triggering svc
with k_sched_lock(). This is unsuitable for Cortex-A/R because:

On SMP, CPUs might need to run their offloaded functions at the same
time (for eg. the smp_abort test) so a single global routine/parameter
pair cannot be used. Also, k_sched_lock() is illegal in interrupt context,
so the shared backend cannot honour CONFIG_IRQ_OFFLOAD_NESTED
(irq_offload() called from an ISR, for eg. test_nested_irq_offload test).

Add a dedicated cortex_a_r/irq_offload.c that fixes the above by:

1. keeping the offloaded routine/parameter per-CPU, indexed by
   _current_cpu->id, so concurrent CPUs no longer interfere with each
   other.

2. using arch_irq_lock()/arch_irq_unlock() instead of k_sched_lock() to
   pin the caller to its CPU while the per-CPU slot is written and the
   svc is taken, while remaining legal from interrupt context. On
   Cortex-A/R the SVC is not masked by CPSR.I. So it still traps with
   interrupts locked.

The offloaded routine/parameter are kept in a per-CPU slot rather than
passed in registers (as arm64 does) to keep this change self-contained
in C and to avoid modifying the shared SVC exception entry used for
context switch, oops and syscalls.

Build the new file for Cortex-A/R, and select ARCH_HAS_IRQ_OFFLOAD_NESTED
for CPU_AARCH32_CORTEX_A and CPU_AARCH32_CORTEX_R.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
9d8ace20d0 arch: arm: cortex_a_r: smp: power on secondary cores via pm_cpu_ops
arch_cpu_start() relied solely on the reset.S voting lock to bring up
secondary cores. That is sufficient only on platforms that release all
cores at reset. On platforms that hold secondaries powered off until
requested (e.g. the Arm FVP Base model), they must be explicitly
powered on with their reset vector pointing at the Zephyr entry point.

Use the generic pm_cpu_on() API to start each secondary at __start.
This dispatches to whichever CPU power driver is enabled. The call is
guarded by CONFIG_PM_CPU_OPS, leaving platforms that release
all cores at reset unaffected.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
f7fb4047c9 arch: arm: cortex_a_r: smp: flush boot params before starting secondary
arm_cpu_boot_params is populated by the primary core and consumed by
the secondary being brought up. On platforms that start secondaries
cold, the secondary begins executing with its caches and MMU disabled
and reads the parameters directly from main memory.

The primary used sys_cache_data_invd_range() on the structure, which
invalidates the cache lines without writing them back. The freshly
written parameters can therefore be discarded before reaching memory,
leaving the secondary to read stale values (e.g. a wrong MPID).

Use sys_cache_data_flush_range() to write the lines back instead,
followed by a DSB to guarantee the flush has completed before the
secondary is brought up.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Silesh C V
5680176ea9 arch: arm: aarch32: mmu: only build shared page tables on primary core
z_arm_mmu_init() runs on every CPU and rebuilds the single, globally
shared L1/L2 page tables on each call. On SMP this is not only
unnecessary duplicate work on the secondary cores (the primary has
already built the tables), but also unsafe. While a secondary core
rebuilds the tables, descriptors that are live and in use by other
cores (the primary or any secondaries with their MMUs enabled) can
temporarily hold invalid values, which can lead to aborts on the
cores using the tables. In practice this manifested as a prefetch
abort on the primary core as soon as a secondary core was brought up
and began (re)building the shared tables.

Fix this by mirroring the arm64 z_arm64_mm_init(is_primary_core)
design: build the tables only on the primary core, in a new
arm_mmu_setup_ptables() helper, and have every core program its MMU
registers and enable translation using the shared tables. As the
page table region's memory attributes are needed while programming
TTBR0, introduce another helper arm_mmu_get_pt_attrs() that is run
by all the cores to look this up.

Also, remove the unused z_arm_mmu_init declaration from
cortex_m/kernel_arch_func.h as that header is included only under
CPU_CORTEX_M.

Single-core behaviour is unchanged.

Signed-off-by: Silesh C V <silesh@alifsemi.com>
2026-09-16 09:22:52 +02:00
Adrian Śliwa
78087d0662 arch: riscv: enable U-mode access to the Zicntr counters
Reads of cycle, time and instret below M-mode are gated by mcounteren
and scounteren, both zero out of reset. Zephyr sets neither in a build
with user threads, so user mode cannot read these counters.

Set them during per-CPU init, since the registers are per-hart. An
M-mode kernel writes mcounteren, and scounteren when misa reports
S-mode, as it does not exist otherwise. An S-mode kernel writes
scounteren alone, because mcounteren is out of reach and is expected
to be set by the SBI, as reset.S does for the built-in one.

Signed-off-by: Adrian Śliwa <asliwa@internships.antmicro.com>
2026-09-16 05:58:39 +02:00
Adrian Śliwa
56164d22a4 arch: riscv: share the counter-enable bit definitions
The counter enable bits are identical in mcounteren, scounteren and
hcounteren, so move them out of reset.S into csr.h. This lets the arch
code that opens the counters to user mode reuse them instead of adding
a second copy.

Signed-off-by: Adrian Śliwa <asliwa@internships.antmicro.com>
2026-09-16 05:58:39 +02:00
Dhruv Menon
908a2e5f0d arch: riscv: switch to main stack before unmasking interrupts
If an interrupt was already pending when z_cstart() completed (such as an
early timer compare event or peripheral interrupt), the ISR fired
immediately while sp still pointed to the interrupt stack. Because
_current_cpu->nested is 0, _isr_wrapper assumed the interruption came
from thread mode and reset sp to the top of _kernel.cpus[0].irq_stack,
causing subsequent ISR C call frames to write directly over the hardware-
saved ESF. This corrupted the saved mepc, leading to an Illegal Instruction
exception (mcause: 2) upon mret.

this commit fixes this by moving interrupt enablement
(csrs RV_STATUS_CSR, %2)  inside the inline assembly block after the stack
pointer has been  switched to main_stack (mv sp, %0)

Signed-off-by: Dhruv Menon <dhruvmenon1104@gmail.com>
2026-09-16 05:55:37 +02:00
Sylvio Alves
886d9899db arch: riscv: add soc hook to close the syscall ecall
Some RISC-V SoCs cannot deliver a fault raised by the syscall
body itself while the ECALL exception of a user-mode syscall is
still being handled. The stack guard below the privileged stack
catches a syscall whose call chain runs too deep, but on such a
SoC that access fault is reported asynchronously inside the open
ECALL and locks the core up, resetting the chip with nothing
delivered to software. A user-mode application can therefore
reset the board through the depth of a granted service call.

Masking interrupts does not help here. The existing
RISCV_SOC_HAS_SYSCALL_INTMASK holds off the interrupt
controller, while this fault comes from the body itself.

Add the hidden RISCV_SOC_SYSCALL_CLOSE_ECALL option. When a SoC
selects it, the IRQ wrapper leaves the exception before running
the syscall body: it returns to the following instruction with
the previous privilege set to machine mode, so the body runs on
the same privileged stack with the same privileges but outside
the exception. A fault taken there is an ordinary top-level trap
the SoC delivers normally, and a syscall that overflows the
privileged stack is reported as a stack overflow that kills only
the offending thread.

The exception exit path reloads the exception program counter
and status register from the saved frame, so borrowing both here
does not disturb the return to user mode.

Assisted-by: Claude:opus-5
Signed-off-by: Sylvio Alves <sylvio.alves@espressif.com>
2026-09-14 15:41:06 -04:00
Sylvio Alves
32381dcbfb arch: riscv: add soc hook to mask syscall interrupts
Some RISC-V SoCs cannot take an interrupt in the syscall body: an
interrupt, or another asynchronously reported event such as an
imprecise access fault, taken while the ECALL exception is still
being handled locks the CPU up, with no fault delivered to
software. Precise synchronous traps, such as the ECALL used to
switch out of the body, nest normally. The syscall path sets
mstatus.MIE before running the body, so such a SoC must hold
interrupts off another way.

This is not the Smdbltrp double trap. On the affected cores
reading mstatush traps, so there is no mstatus.MDT to clear, and
the vendor status CSR exposes the in-exception state as a
read-only bit that stays set until mret. Software cannot end the
exception window early, so the only option is to keep interrupts
from being taken while it is open.

Add the hidden RISCV_SOC_HAS_SYSCALL_INTMASK option following
the existing RISCV_SOC_HAS_* hook pattern. When a SoC selects it,
the IRQ wrapper invokes a SoC-provided SOC_SYSCALL_INTMASK macro
on syscall entry, typically raising a hardware interrupt level
threshold, while the generic code carries no SoC register
knowledge. This is a SoC hook rather than part of the generic
CLIC interrupt level support because the SoCs that need it drive
their interrupt controller from SoC code and do not select
RISCV_HAS_CLIC.

The mask is per exception frame. The SoC context hooks save the
state on entry and restore it on the exception exit path, and
with RISCV_ALWAYS_SWITCH_THROUGH_ECALL every context switch goes
through that path: a thread that blocks in the body hands the
CPU to a thread that restores its own saved state, and a new
thread starts from SOC_ESF_INIT. The option therefore depends on
both. While the mask is raised the body is not preemptible and
pending interrupts, the tick included, wait for the syscall to
return or block; the help text says so.

Assisted-by: Claude:opus-5
Signed-off-by: Sylvio Alves <sylvio.alves@espressif.com>
2026-09-14 15:41:06 -04:00
Sylvio Alves
37d3022714 arch: riscv: pre-validate user strings on imprecise-fault socs
arch_user_string_nlen() dereferences a user pointer and recovers
from a bad address through a fault fixup that matches the
faulting mepc against the load instruction's range. On some
RISC-V SoCs the load access fault is imprecise: the reported mepc
lands past the faulting load, so the fixup misses it and the
fault escalates to a fatal reset.

Rename the asm routine to z_riscv_user_string_nlen() and alias
arch_user_string_nlen() to it by default. Under the new
RISCV_USER_STRING_NLEN_VALIDATE option a C implementation takes
over instead: it checks the whole range with
arch_buffer_validate() first and, when that fails, walks the
string in PMP-granularity chunks, validating each chunk before
reading it, so no load that could fault is ever issued. The
faulting load may sit at any offset (a string that starts in an
accessible region and runs unterminated into an inaccessible
one), which is why one check of the first byte is not enough.
The generic syscall handler is untouched.

Add a syscalls test scenario that enables the option on RISC-V.

Assisted-by: Claude:opus-5
Signed-off-by: Sylvio Alves <sylvio.alves@espressif.com>
2026-09-14 15:41:06 -04:00
Benjamin Cabé
a811b0f1cd arch: arm: cortex_m: drop redundant ISBs from IRQ lock and unlock
arch_irq_lock() and arch_irq_unlock() executed an ISB after every
BASEPRI write on ARMv7-M/ARMv8-M Mainline, and after CPSIE on ARMv6-M/
ARMv8-M Baseline unlock: two pipeline flushes for every kernel critical
section. Neither is architecturally required. Priority-raising MSR
writes are self-synchronizing (Arm DAI 0321A, section 4.8), and on
unlock a pended interrupt is taken within a couple of instructions
anyway; only arch_swap() relies on the pended PendSV being recognized
before the next instruction executes, so it gains an explicit ISB.

Cortex-M7 r0p0/r0p1 keeps the ISB on the raising side, where erratum
440977 (formerly 837070) can delay a priority-raising BASEPRI write. The
CONFIG_CORTEX_M_ERRATUM_440977_WORKAROUND option defaults to y on M7;
the same ISB in the PendSV handler prologue is handled identically.

The ISB occupies the instruction that can still be preempted before the
critical region begins. Retain this existing sequence to avoid masking
higher-priority interrupts through PRIMASK, as Arm's documented
CPSID i/MSR/CPSIE i workaround would do.

Measured standalone against main (latency_measure in cycles, lower is
better; thread_metric scores, higher is better):

* az3166_iotdevkit (STM32F412, Cortex-M4 @ 96 MHz):
  - latency_measure: mean -3.3% over 47 ops, min/max/median
    -8.1/+0.6/-2.9%; semaphore give no-waiter 74 -> 68 cycles.
  - thread_metric: synchronization +15.4%, cooperative +3.3%,
    preemptive +4.6%, interrupt +2.9%.

* mps2/an385 (Cortex-M3, QEMU icount):
  - latency_measure: mean -3.8%, min/max/median -6.8/+0.0/-3.8%.
  - thread_metric: synchronization +8.3%.

* Flash, az3166 latency_measure image: -400 B at -Os, -424 B at -O2.

Assisted-by: Codex:gpt-6-astra
Signed-off-by: Benjamin Cabé <benjamin@zephyrproject.org>
2026-09-14 08:42:05 -04:00
Hongquan Li
dc577c7239 arch: arm64: fix DAIF restore on arch_buffer_validate() wrap path
Preserve the DAIF state across failed buffer validation.

Fixes #115141

Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
2026-09-13 21:38:39 -04:00
Andrei-Edward Popa
ef5b9dc5c1 arch: arm: mpu: fix region init for Cortex-R
ARM_MPU_REGION_INIT() currently includes the MPU region size in the region
attributes for both Cortex-M and Cortex-R.

On Cortex-M, the region attributes and size are both programmed through
RASR, so this is correct. Cortex-R uses separate registers for region
access attributes and region size. Including the size in the attributes
therefore corrupts the memory type attributes programmed into DRACR.

Store the encoded region size in the dedicated size field for Cortex-R
and exclude it from the region attributes. This fixes devicetree-defined
MPU regions on Cortex-R, where the region size could alter the TEX, S, C,
and B attribute bits.

Signed-off-by: Andrei-Edward Popa <andrei.popa105@yahoo.com>
2026-09-13 21:38:15 -04:00
Hongquan Li
a0824279b9 arch: arm64: require stack protection for safe exception stack
ARM64_SAFE_EXCEPTION_STACK is user-selectable without
ARM64_STACK_PROTECTION, but thread->arch.stack_limit is only assigned
when the latter is on (arch_new_thread). With the safe exception stack
enabled standalone, current_stack_limit stays 0, so the overflow check
in z_arm64_quick_stack_check() can never trip: every EL1 exception pays
the entry overhead and the kernel stack overflow detection silently
does not exist.

Make ARM64_STACK_PROTECTION the only way in. No in-tree configuration
is affected: no AArch64 target selects ARM_MPU today, so
ARM64_SAFE_EXCEPTION_STACK was never enabled in any built config.

Fixes #118385

Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
Assisted-by: Codex:GPT-5 gh
2026-09-13 21:38:07 -04:00
Daniel Leung
2756f2c863 xtensa: mmu: simplify TLB shootdown
Since we are caching all the necessary register values to be
used when switching page tables, we can now remove the extra
bits inside xtensa_mmu_tlb_shootdown() originally used to
prevent re-computing all those register values. We can now
simply call xtensa_mmu_set_paging().

Signed-off-by: Daniel Leung <daniel.leung@intel.com>
2026-09-13 17:59:42 -04:00
Daniel Leung
60c9284620 xtensa: mmu: no TLB shootdown if not needed during thread add
Inside arch_mem_domain_thread_add(), if the thread is not
running on other CPUs and it is not migrating between domains,
there is no need to send TLB shootdown to other CPUs. Next time
the thread is scheduled, the new page tables will be used.

Signed-off-by: Daniel Leung <daniel.leung@intel.com>
2026-09-13 17:59:42 -04:00
Brian Cain
8b64665e96 hexagon: implement arch_irq_offload()
Run a routine in interrupt context by posting a software interrupt
through the VM and letting it fire when irq_unlock() re-enables IE.

IRQ 0 is used as the trigger line: the board's DTS assigns the lowest
hardware IRQ at 0x0c, so it does not collide with a real device here.

Give the test suite a trigger_irq() for Hexagon as well, going through
the same hypercall, so the shared interrupt tests can drive a line from
software.

Signed-off-by: Brian Cain <brian.cain@oss.qualcomm.com>
2026-09-12 07:52:39 -04:00
Brian Cain
5761d7d707 hexagon: add architecture port
Add Hexagon DSP architecture support to Zephyr RTOS, targeting the
Qualcomm Hexagon V67+ ISA running as a guest under the H2 hypervisor.

This includes:
- Architecture scaffolding (Kconfig, CMakeLists, arch selection)
- Public headers (arch.h, thread.h, exception.h, error.h)
- Core runtime: context switch, interrupt/exception handling, idle, TLS
- Memory management stubs for flat-memory H2 guest model
- SoC support for QEMU hexagon virt machine
- CMake/LLVM toolchain integration for hexagon cross-compilation

Hexagon uses H2's GEVB (Guest Event Vector Base) for interrupt dispatch
rather than a traditional function-pointer vector table, so
GEN_IRQ_VECTOR_TABLE is disabled by default.

Signed-off-by: Brian Cain <brian.cain@oss.qualcomm.com>
2026-09-12 07:52:39 -04:00
Liu Qian
18b49e8379 arch: riscv: pmp: add NAPOT multi-slot mode for non-aligned regions
When TOR is unsupported (PMP_NO_TOR), the Kconfig forced
PMP_POWER_OF_TWO_ALIGNMENT=y, wasting memory: 33 KB of code+rodata
aligns to 64 KB, 257 KB aligns to 512 KB.

Add PMP_NAPOT_USE_MULTI_SLOTS: splits a non-naturally-aligned region
into multiple NAPOT entries, each covering the largest naturally-
aligned block at the current address. [0x1800, 0x3800) becomes
0x800 + 0x1000 + 0x800 (3 slots instead of padding to 8192).

Depends on !USERSPACE: arch_mem_domain_max_partitions_get() assumes
1-2 slots per partition and resync_pmp_domain() can only log on
failure, so a partition needing more slots would run unmapped (silent
loss of protection). Help text documents that disabling
PMP_POWER_OF_TWO_ALIGNMENT also splits stack guards and u-mode stacks,
and that PMP_GRANULARITY=64 cuts both ways for that case.

try_multi_entries_set() uses clz/ctz builtins, clears already-written
config bytes on failure before restoring *index_p (otherwise the
trailing write_pmp_entries(0, PMP_SLOTS) in z_riscv_pmp_init() would
program locked partial entries), and folds start/size into the inner
failure logs.

Signed-off-by: Liu Qian <liuqian.andy@picoheart.com>
2026-09-11 17:34:41 +02:00
Liu Qian
96c7e335e4 arch/riscv: add release barrier before publishing switch_handle
Under RVWMO the context stores (callee-saved registers, stack
pointer) may be reordered before the switch_handle store, so a hart
spinning in z_sched_switch_spin() could observe the handle before
the saved context is visible and restore stale registers.

Add fence rw, w before publishing switch_handle, pairing with the
acquire barrier in z_sched_switch_spin(), same as ARM64.

Signed-off-by: Liu Qian <liuqian.andy@picoheart.com>
2026-09-11 17:32:12 +02:00