Use the existing custom-current architecture hooks to keep the running
thread pointer in TPIDR_EL1. A register read avoids the call and IRQ
masking needed to prevent migration between a CPU lookup and the
current-thread load.
Keep the per-CPU current pointer for cross-CPU readers. The scheduler
updates the register with preemption disabled, and reset initializes it
before the kernel installs the dummy thread on each CPU. TPIDR_EL0
remains available for TLS and TPIDRRO_EL0 for CPU and exception state.
Enable this for Cortex-A SMP without system power management.
Uniprocessor and PM builds keep the existing lookup; power-state
restoration is not extended to save the register.
A53 quad core microbenchmarks show ~1.2% higher throughput and up
to 4.7% lower latency.
Assisted-by: Codex:GPT-5
Signed-off-by: Aaron Wisner <aaronwisner@gmail.com>
ARM64_SAFE_EXCEPTION_STACK is user-selectable without
ARM64_STACK_PROTECTION, but thread->arch.stack_limit is only assigned
when the latter is on (arch_new_thread). With the safe exception stack
enabled standalone, current_stack_limit stays 0, so the overflow check
in z_arm64_quick_stack_check() can never trip: every EL1 exception pays
the entry overhead and the kernel stack overflow detection silently
does not exist.
Make ARM64_STACK_PROTECTION the only way in. No in-tree configuration
is affected: no AArch64 target selects ARM_MPU today, so
ARM64_SAFE_EXCEPTION_STACK was never enabled in any built config.
Fixes#118385
Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
Assisted-by: Codex:GPT-5 gh
A thread actively running on another CPU at panic time was previously
invisible to real debugging: coredump only captured the panicking CPU's
exception frame, so every other thread's registers came from its
k_thread.callee_saved struct -- correct for genuinely sleeping threads,
but stale (reflecting whenever it last voluntarily context switched) for
one that's actually running elsewhere right now.
Add CONFIG_DEBUG_COREDUMP_SMP_FREEZE_CPUS (default y where the arch sets
ARCH_SUPPORTS_COREDUMP_SMP_FREEZE, i.e. arm64 Cortex-A SMP). On panic,
coredump() freezes every other online CPU before dumping and thaws them
after the memory-region walk, and emits each frozen CPU's live registers
as a new coredump section (COREDUMP_CPU_SNAPSHOT_HDR_ID, tagged with CPU
index and thread pointer, reusing the arch-info block register layout).
The panicking CPU (never frozen, since freeze only targets other CPUs)
emits a zero-payload marker recording only which thread was panicking
there, which the host side uses to identify it unambiguously.
arm64 implementation (arch/arm64/core/smp.c, coredump.c): send a new SGI
to every other online CPU. Each handler captures its exact live state --
x0-x18/lr/spsr/elr via the arch_esf that _isr_wrapper() stashes at a
fixed, nesting-independent slot on its IRQ stack, plus x19-x29 via inline
asm (valid per AAPCS64) -- then spins holding whatever it held before
being frozen, untouched, until the dump finishes, and resumes via the
standard exception-return path exactly where it was interrupted. This
matters because a panic here can be recoverable (this tree's own
coredump_threads test doesn't halt/reboot after dumping).
The SGI uses a deliberately low priority (SGI_COREDUMP_FREEZE_PRIO), not
IRQ_DEFAULT_PRIORITY: a frozen CPU never returns from its handler so
_isr_wrapper() never EOIs, and per GIC priority rules only a strictly
higher-priority interrupt can preempt a still-active one. At the default
priority a frozen CPU could not service same-priority peripheral
interrupts (confirmed on hardware as Ethernet TX-done starvation during
the UDP backend); the lowest usable priority lets everything else keep
making progress while a CPU waits.
The freeze/thaw handshake uses two separate per-CPU signals: freeze_state
(IDLE/REQUESTED/CAPTURED) and a dedicated release_requested flag that
only thaw writes. Kept separate so a late-arriving CPU (whose SGI was
delayed past the freeze-side timeout) still sees a release signal instead
of stomping it with its own CAPTURED update and spinning forever; thaw
sets release_requested unconditionally for every other CPU. The handler
wait is also bounded by a 60-second wall-clock backstop (k_cycle_get_32(),
safe from any context) so no CPU is ever stuck indefinitely. A CPU that
never responds within the freeze-side timeout (never booted, or stuck
with IRQs masked) is skipped and never blocks the dump.
The host coredump log parser learns to parse the new snapshot section.
Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
Select ARCH_SUPPORTS_COREDUMP_THREADS and ARCH_SUPPORTS_COREDUMP_STACK_PTR
for CPU_CORTEX_A in arch/arm64/core/Kconfig, enabling MEMORY_DUMP_THREADS
coredump mode on ARM64. Mirrors the existing Cortex-M declarations in
arch/Kconfig.
Implement arch_coredump_stack_ptr_get() using thread->callee_saved.sp_elx,
the EL1 stack pointer saved by the context-switch path for sleeping
threads. For the faulting thread the saved context is stale, so the
coredump falls back to dumping the full stack allocation,
which is always correct.
Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
PRIVILEGED_STACK_SIZE already has ARM64 defaults earlier in the
same Kconfig file.
The later unconditional default is never reached because the previous
unconditional default is the first matching default when FPU_SHARING is
disabled.
Remove the duplicate entry to avoid implying that the value is
overridden later.
Signed-off-by: Hongquan Li <hongquan.li@processmission.com>
Add an AArch64 PMUv3 implementation behind CONFIG_ARM64_PMUV3 (pmuv3.c):
probe ID_AA64DFR0_EL1, calibrate CPU frequency (PMCCNTR_EL0 vs the
generic timer), and provide per-CPU counter configuration, enable/disable,
overflow handling, and cycle counter access. Initialization is explicit
via pmu_init() on each logical CPU that uses the PMU (no SYS_INIT).
Introduce include/zephyr/pmu.h for the portable pmu_*() API.
Architectural PMUv3 event codes (PMU_EVT_* in 0x00-0x1F) and
PMCR/PMUSERENR bit defines live in include/zephyr/arch/arm64/pmuv3.h for
AArch64 builds.
Add ARCH_HAS_PMU in arch/Kconfig (Cortex-A profiles select it); enable
CONFIG_ARM64_PMUV3 for the PMUv3 driver backend. Register access uses
explicit MRS/MSR inlines instead of read_sysreg()/write_sysreg()
statement expressions for static analysis.
Builds for versal_apu, versalnet_apu, and versal2_apu. On QEMU, PMU
access is often unavailable (-ENOTSUP). On Versal Net APU hardware with
PMU usable at the current EL, initialization succeeds.
Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
Tuning CONFIG_MAX_XLAT_TABLES is currently trial-and-error: the only
feedback on overflow is a "too small" panic with no hint of how much
to bump.
Track the high-water mark of allocated translation tables and use it
to emit two signals from new_table(): a one-shot LOG_WRN when the
pool drops below 12.5 % free (always compiled in, advance notice
before the panic), and an opt-in LOG_INF on every new peak (under
CONFIG_ARM64_MMU_REPORT_XLAT_TABLES_USAGE) so the last logged "peak
N of M allocated" line gives a concrete lower bound on MAX for the
workload that ran.
Signed-off-by: Carlo Caione <ccaione@baylibre.com>
Add new implementations for entropy driver and random subsystem
based on ARM64 RNDRRS and RNDR instructions.
Signed-off-by: Christoph Busold <cbusold@qti.qualcomm.com>
Modern ARM64 SoCs with SMP and 36-bit virtual addressing require more
translation tables than the current default of 8. This is particularly
evident on TI K3 SoCs (AM62X, AM62LX) where there are more SoC
peripherals which are beyond the 2MB boundary.
Add a new default: 16 tables for SMP && (ARM64_VA_BITS >= 36)
Signed-off-by: Soumya Tripathy <s-tripathy@ti.com>
Since commit 0026a5610ac ("arm64: mm: use identity mapping for device
MMIO"), device_map() creates identity mappings (VA = PA) instead of
allocating virtual addresses from a contiguous pool. Each device at a
distinct 2MB-aligned physical address now requires its own L3 page
table, increasing the total number of translation tables needed.
Bump the USERSPACE && TEST default from 24 to 28 to accommodate the
additional tables required by identity-mapped device MMIO.
Signed-off-by: Carlo Caione <ccaione@baylibre.com>
On ARM64, Zephyr uses identity mappings (VA = PA) for kernel code, data and
boot-time device regions. The MMU fully supports address translation but
Zephyr uses it primarily for access permission enforcement.
There are currently two independent paths for mapping device MMIO regions:
1. SoC-level mmu_regions.c files use MMU_REGION_FLAT_ENTRY() to create
identity mappings (VA = PA) directly in the page tables at boot. This
bypasses the kernel's virtual memory tracking entirely. SoC maintainers
must manually list peripherals in mmu_regions.c for drivers that do not
use the device MMIO API (e.g. most existing drivers) or cannot use it
(e.g. the GIC, which is not a regular driver).
2. The device MMIO API (device_map()) goes through k_mem_map_phys_bare(),
which allocates a virtual address from the SRAM range and maps (VA !=
PA) device registers there. Mapping device MMIO into the SRAM virtual
address space is nonsensical: it conflates device registers with memory,
wastes virtual address pool space, and produces addresses that bear no
relation to the hardware.
The CONFIG_KERNEL_DIRECT_MAP mechanism already supports identity mapping
through k_mem_map_phys_bare() with the K_MEM_DIRECT_MAP flag, but it
requires each board defconfig to enable the Kconfig and each driver to
explicitly pass the flag.
Make identity-mapped device MMIO automatic on ARM64:
1. ARM64 CPU_CORTEX_A selects KERNEL_DIRECT_MAP when MMU is enabled. This
eliminates the need for per-board defconfig opt-in.
2. device_map() automatically injects K_MEM_DIRECT_MAP when
CONFIG_KERNEL_DIRECT_MAP is enabled. This is transparent to
drivers so no per-driver changes needed. The flag is gated on
CONFIG_KERNEL_DIRECT_MAP rather than CONFIG_ARM64, keeping it
architecture-agnostic.
Signed-off-by: Carlo Caione <ccaione@baylibre.com>
Implement thread-based unwinding to support the 'kernel thread unwind'
shell command on ARM64. This update ensures that 'thread' defaults to
'_current' when NULL, complying with the arch_stack_walk() API contract.
To enhance security, add stack bounds validation using stack_info,
and TLS pointers of the target thread. If these are not available, the
logic correctly falls back to is_address_mapped() to ensure robustness
during the unwinding process.
Signed-off-by: Archilis Wang <awm02289@gmail.com>
Add ARM64_PAGE_SIZE Kconfig choice allowing 4KB, 16KB and 64KB
page sizes. The MMU code already derived all constants from
PAGE_SIZE_SHIFT so most of the infrastructure was ready.
Changes:
- Add ARM64_PAGE_SIZE choice (4KB default, 16KB, 64KB) in Kconfig
- Derive PAGE_SIZE_SHIFT from CONFIG_MMU_PAGE_SIZE in mmu.h
- Select proper TCR granule bits (TG0/TG1) per page size in mmu.c
- Round ARCH_THREAD_STACK_RESERVED up to page alignment so that
the user-accessible stack buffer starts on a page boundary
- Fix MEM_REGION_ALLOC in mem_protect test to use CONFIG_MMU_PAGE_SIZE
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Implement arch_mem_domain_deinit() for ARM64 to release page tables
back to the pool when a memory domain is de-initialized. This reuses
the existing discard_table() mechanism to recursively free all
sub-tables in the hierarchy.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
After commit 02770ad963 ("debug: EXCEPTION_STACK_TRACE should depend
on arch Kconfigs"), the ARM64_EXCEPTION_STACK_TRACE isn't used any more,
remove it.
Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
Move ARCH_HAS_STACKWALK under CPU_CORTEX_A section since only Cortex-A
implements arch_stack_walk(), while Cortex-R does not.
Signed-off-by: Sudan Landge <sudan.landge@arm.com>
Memory protection and userspace tests require more MMU translation
tables than the default. Without this increase, tests fail with:
E: CONFIG_MAX_XLAT_TABLES too small
ASSERTION FAIL [ret == 0] @ arch/arm64/core/mmu.c:1244
privatize_page_range() returned -12
Increase defaults when both USERSPACE and TEST are enabled:
- 32 tables for SMP configurations
- 24 tables for non-SMP configurations
This fixes:
- sample.kernel.memory_protection.shared_mem (all platforms)
- rtio.api.userspace (v8a, v9a)
- rtio.api.userspace.submit_sem (v8a, v9a)
- portability.posix.common.userspace
Consequently the demand paging test needed adjustment to its
qemu_cortex_a53 configs to keep working as this test is highly
sensitive to the amount of available free memory.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Increase ARM64 stack sizes to accommodate deeper call stacks in
userspace and SMP configurations when FPU_SHARING is enabled:
- PRIVILEGED_STACK_SIZE: 1024 → 4096 bytes (with FPU_SHARING)
- TEST_EXTRA_STACK_SIZE: 2048 → 4096 bytes (with FPU_SHARING)
The default 1KB privileged stack is insufficient for ARM64 userspace
syscalls when FPU context switching is enabled.
Symptom: Userspace tests crash with Data Abort (EC 0x24) near stack
boundaries during syscalls, particularly on SMP configurations where
multiple threads exercise FPU lazy switching.
Fixes previously failing CI test on fvp_base_revc_2xaem SMP variants:
- kernel.threads.dynamic
- Multiple userspace tests with FPU_SHARING enabled
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Implement Scalable Vector Extension (SVE) context switching support,
enabling threads to use SVE and SVE2 instructions with lazy context
preservation across task switches.
The implementation is incremental: if only FPU instructions are used
then only the NEON access is granted and preserved to minimize context
switching overhead. If SVE is used then the NEON context is upgraded to
SVE and then full SVE access is granted and preserved from that point
onwards.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Add Cortex-A320 support to the unified FVP board structure with ARMv9.2-A
specific configuration parameters.
New board target:
- fvp_base_revc_2xaem/a320
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Add ARMv9-A architecture support with Cortex-A510 CPU as the default
processor for generic ARMv9-A targets.
Signed-off-by: Nicolas Pitre <npitre@baylibre.com>
Disabling multithreading is not possible when enabling SMP (logically)
so depend on SMP being disabled to enable
ARCH_HAS_SINGLE_THREAD_SUPPORT.
Signed-off-by: Carles Cufi <carles.cufi@nordicsemi.no>
Introduce a new Kconfig option CPU_CORTEX_A78 to enable support for the
Arm Cortex-A78 CPU architecture within Zephyr. This configuration can be
selected by boards or SoCs that utilize the Cortex-A78 core, enabling
architecture-specific features and optimizations as needed.
Signed-off-by: Appana Durga Kedareswara rao <appana.durga.kedareswara.rao@amd.com>
In the commit 573a712bed patch "arm64:
reset: disable cache and MMU for safety", it disables D-Cache and MMU
for safety, but in some cases, for example the code is loaded into memory
by hardware debugger, we need to flush D-Cache before disable it in
order to make sure the data is coherent in the system, otherwise it
will report "Synchronous Abort" when D-Cache is disabled.
Signed-off-by: Jiafei Pan <Jiafei.Pan@nxp.com>
Added new configuration item to optionally enable APIs for operation
all data cache, by default these APIs are disabled.
Signed-off-by: Jiafei Pan <Jiafei.Pan@nxp.com>
GCC and Clang support the undefined behavior sanitizer in any
configuration, the only restriction is that if you want to get nice
messages printed, then you need the ubsan library routines which are only
present for posix architecture or when using picolibc.
This patch adds three new compiler properties:
* sanitizer_undefined. Enables the undefined behavior sanitizer.
* sanitizer_undefined_library. Calls ubsan library routines on fault.
* sanitizer_undefined_trap. Invokes __builtin_trap() on fault.
Overhead for using the trapping sanitizer is fairly low and should be
considered for use in CI once all of the undefined behavior faults in
Zephyr are fixed.
Signed-off-by: Keith Packard <keithp@keithp.com>
`CONFIG_ARM64_ENABLE_FRAME_POINTER` had been deprecated since #72646
for 2 releases and served not functional effect, it's now time to
say goodbye.
Signed-off-by: Yong Cong Sin <ycsin@meta.com>
Signed-off-by: Yong Cong Sin <yongcong.sin@gmail.com>
Currently it supports `esf` based unwinding only.
Then, update the exception stack unwinding to use
`arch_stack_walk()`, and update the Kconfigs & testcase
accordingly.
Also, `EXCEPTION_STACK_TRACE_MAX_FRAMES` is unused and
made redundant after this change, so remove it.
Signed-off-by: Yong Cong Sin <ycsin@meta.com>
Signed-off-by: Yong Cong Sin <yongcong.sin@gmail.com>
Currently, the stack trace in ARM64 implementation depends on
frame pointer Kconfigs combo to be enabled. Create a dedicated
Kconfig for that instead, so that it is consistent with x86 and
riscv, and update the source accordingly.
Signed-off-by: Yong Cong Sin <ycsin@meta.com>
Introduce the ARM64_STACK_PROTECTION config. This option leverages the
MMU or MPU to cause a system fatal error if the bounds of the current
process stack are overflowed. This is done by preceding all stack areas
with a fixed guard region. The config depends on MPU for now since MMU
stack protection is not ready.
Signed-off-by: Jaxson Han <jaxson.han@arm.com>
Xen-related Kconfig options were highly dependand on BOARD/SOC xenvm.
It is not correct because Xen support may be used on any board and SoC.
So, Kconfig structure was refactored, now CONFIG_XEN is located in
arch/ directory (same as in Linux kernel) and can be selected for
any Cortex-A arm64 setup (no other platforms are currently supported).
Also remove confusion in Domain 0 naming: Domain-0, initial domain,
Dom0, privileged domain etc. Now all options related to Xen Domain 0
will be controlled by CONFIG_XEN_DOM0.
Signed-off-by: Dmytro Firsov <dmytro_firsov@epam.com>
Enhanced arch_start_cpu so if a core is not available based on pm_cpu_on
return value, booting does not halt. Instead the next core in
cpu_node_list will be tried. If the number of CPU nodes described in the
device tree is greater than CONFIG_MP_MAX_NUM_CPUS then the extra cores
will be reserved and used if any previous cores in the cpu_node_list fail
to power on. If the number of cores described in the device tree matches
CONFIG_MP_MAX_NUM_CPUS then no cores are in reserve and booting will
behave as previous, it will halt.
Signed-off-by: Chad Karaginides <quic_chadk@quicinc.com>
Introduce two configs to prepare to enable the safe exception stack for
the kernel space. This is the preparation for enabling hardware stack
guard. Also define the safe exception stack for kernel exception stack
check.
Signed-off-by: Jaxson Han <jaxson.han@arm.com>
VMPIDR_EL2 is assigned the value returned by EL2 reads of MPIDR_EL1
MPIDR_EL1 is the register holding the Multiprocessor ID which is to
identify different cores. Because of the virtualization requirements
for AArch64, MPIDR_EL1 should be virtualized (the different virtualized
cores can run on the same physical core). Thus the value of MPIDR_EL1
should be switched when the VM is switched. Setting the VMPIDR_EL2 is
the way to change the value returned by EL1 reads of MPIDR_EL1. Even
without virtualization, we still need to set VMPIDR_EL2 during booting
at EL2 or EL3. Otherwise, all cores' IDs are zero at the EL1 stage
which will break the SMP system.
Signed-off-by: Huifeng Zhang <Huifeng.Zhang@arm.com>
Enable single-threaded support for the arm64 archtecture.
This mode of execution is supported on an soc under
development and is validated regularly.
Signed-off-by: Eugene Cohen <quic_egmc@quicinc.com>
On platforms where reset vector catch is not possible
it is useful to have a compile-time option to spin
at the reset vector allowing a debugger to be attached
and then to manually resume execution.
Define a config option for arm64 to spin at the
reset vectdor so a debugger can be attached.
Signed-off-by: Eugene Cohen <quic_egmc@quicinc.com>
In some drivers, noncache memory need to be used for dma coherent
memroy, so add nocache memory segment mapping and support for ARM64
platforms.
The following variables definition example shows they will use nocache
memory allocation:
int var1 __nocache;
int var2 __attribute__((__section__(".nocache")));
Signed-off-by: Jiafei Pan <Jiafei.Pan@nxp.com>