Rethinking the Stack: DDR-Free, Zephyr-First Design on SoC FPGAs
Cut complex, power-hungry DDR. A Zephyr-first design on PolarFire SoC FPGAs leverages on-chip scratchpad and low-power alternatives to boost efficiency and lower BoM costs.
Most SoC designs follow the same system architecture: Linux running on the top of the stack, using DDR as your main memory. It is mature, well supported, well understood and great for feature-rich applications. But it carries assumptions worth questioning, particularly for applications where real-time behaviour, power efficiency and BoM cost matter more than OS feature richness.
What if Linux wasn’t your default starting point and Zephyr was instead? A Linux-first approach can come with more requirements than Zephyr-first, depending on what you want to do. If you care about determinism , power efficiency, BoM cost or all of the above, questioning the conventional approach could help you meet system goals.

An SMP Zephyr architecture using the PolarFire SoC FPGA.
An alternative system architecture is worth considering: SMP Zephyr at the top of the stack, running on all application cores, executing entirely from on-chip memory and eliminating external DDR from the platform. This model rethinks the hardware–software partitioning and memory architecture from the ground up, favoring tightly coupled compute, predictability and simplified system design over traditional OS-centric approaches.
Before We Talk About Removing DDR—How Much Bandwidth Do You Actually Need?
Let's start with an honest question that most designs skip: what is your application actually doing with memory bandwidth?
DDR exists because Linux needs it. A modern Linux SoC running a full OS stack, kernel, userspace, drivers, file systems, generates memory traffic that only DDR can handle. That is a legitimate requirement. But it is Linux's requirement, not necessarily yours.
If you are running a control loop, a signal processing pipeline, a communications stack or a set of deterministic real-time tasks, your actual bandwidth budget might be far smaller than you assume. Before committing to the cost and complexity of DDR, it is worth doing that calculation explicitly rather than inheriting it from the OS choice.
And if the answer is "some external memory, but not DDR-scale bandwidth," you have more options than most engineers realise.
The Memory Spectrum Nobody Talks About
The embedded memory landscape between "on-chip SRAM" and "full DDR" is larger and more capable than it gets credit for. Two options worth serious consideration:
HyperRAM connects over a simple 12-pin HyperBus interface, requires no complex termination or length-matched routing and delivers bandwidth in the range of 333 MB/s—more than sufficient for many embedded workloads that would otherwise reach for DDR. The PCB complexity difference is dramatic. No DDR PHY, no matched pairs, no signal integrity review cycle. A junior engineer can route it.
IoT PSRAM takes this further. Low-pin-count SPI or QPI interfaces, ultra-low power and enough bandwidth for applications that need a working set larger than on-chip memory but are nowhere near saturating DDR. These parts are designed specifically for the class of application where DDR is overkill but on-chip alone is tight.

Memory choices based on workload.
The point is not that these replace DDR for every use case. The point is that the moment you let go of the assumption that it has to be DDR, a whole tier of design choices opens up. Choices that trade raw peak bandwidth for dramatically lower cost, simpler layout, lower power and faster time to a working board.
Most MPU-based systems are locked to a fixed memory hierarchy: your options are narrow and they almost always lead back to DDR. SoC FPGAs change that. PolarFire SoC opens up the design space, letting you rethink how the memory subsystem is built rather than inheriting someone else's assumptions.
Attaching a HyperRAM or PSRAM device is a natural fit. The fabric handles the interface logic; the processor complex sees clean, memory-mapped resources. You get expanded address space without the cost, power and complexity that come with the DDR tax.
Why DDR Is Worth Cutting (When You Can)
When DDR is the right call, use it. But when the bandwidth analysis says it is not required, cutting it removes a surprising amount of design overhead.
Bill of materials. DDR components carry non-trivial cost and supply chain volatility has made lead time and sourcing risk a real concern in recent years.
PCB complexity. DDR routing is among the most demanding signal integrity work in any board design—matched-length differential pairs, controlled impedance, careful power delivery and the layout review cycles that come with it. That is engineering time and cost that does not appear on the component line of a BOM.
Power consumption. DDR draws meaningful idle power and the power delivery infrastructure it requires adds overhead before your application has done any useful work.
Latency and non-determinism. DDR refresh cycles, bank conflicts and memory controller overhead all introduce variability that is difficult to eliminate from a real-time system, no matter how carefully you design around it.
The DDR-free path—whether that means pure on-chip scratchpad or scratchpad plus a HyperRAM or PSRAM for overflow—collapses all of this overhead and gives it back to you as design margin.
What the Scratchpad Actually Gives You and How Far Zephyr Fits Inside It
PolarFire SoC includes tightly-coupled on-chip scratchpad memory providing at least 1.5 MB of usable code and data space. Accessed directly by the processor without going through a memory controller, this delivers consistent, low-latency access. Exactly what a deterministic real-time workload needs.
The reason this matters so much for Zephyr specifically is the kernel's footprint. A minimal Zephyr build targeting PolarFire SoC comes in at around 20 KB. That is not a typo. The Zephyr kernel is genuinely lean. It includes only what you configure in, with no hidden OS overhead quietly consuming memory in the background.
From that 20 KB baseline, the image grows with what you add. A realistic application with drivers, a thread pool, a networking stack and meaningful application logic will land well below the scratchpad ceiling. To illustrate what Zephyr can carry while staying entirely on-chip, consider a complete edge AI workload: a single-chip parking lot camera that counts empty parking spaces on the device itself and reports the result back to base, where it can be shown on a website or your maps app and you know if you should park in a lot.
The device will capture a frame every second or so, see how many free spaces are there and report it back. No sending of images back to base or processing required back home. It uses a camera (for simplicity on SPI) to capture images, a JPEG decoder, an int8 convolutional neural network running under TensorFlow Lite Micro, a full IPv4/TCP networking stack over a 10BASE-T1S link, and an encrypted MQTT telemetry uplink secured with TLS 1.2, certificate parsing and all.

Usable on-chip scratchpad, with and without encryption.
The whole thing lands at roughly 850 KB comfortably under 1 MB, a little over half the scratchpad with no external memory of any kind. Strip the encryption layer back out and the same vision-plus-networking node sits just over 500 KB. The instructive part is what that buys you: a feature-complete AI vision pipeline, engineered around small compressed buffers rather than full-resolution rasters, keeps both its code and its data inside a scratchpad—no memory controller, no external DRAM, no flash to attach.
And the chip running it could equally have booted Linux, which would have meant populating external DDR, attaching Flash for a root filesystem and waiting tens of seconds to start, where the Zephyr node needs none of that and is alive in well under a second.
For the majority of embedded applications, industrial control, communications, motor drive, sensor fusion and increasingly on-device AI, the entire application stack stays on-chip. From a 20 KB kernel baseline to a full-featured, secure edge AI node at around 850 KB, with 1.5 MB available. The arithmetic is comfortable.
Don’t forget about the FPGA though. That can unlock even more functionality that you might not get in a typical ASIC, we’ll come back to it later.

Microchip PolarFire SoC FPGA with Zephyr SMP on 4 U54 application cores.
Zephyr SMP: Four Cores, Real Parallelism
Zephyr RTOS has matured significantly in recent years, and its SMP support is one of the most important developments for multicore application processors. On PolarFire SoC's quad-core U54 complex, this means:
- Threads genuinely run in parallel across cores, not time-sliced, but truly simultaneous
- Core affinity and pinning are supported, giving you deterministic placement of critical workloads
- The scheduler is real-time throughout, no Linux-style background housekeeping competing with your application
This isn't an argument that Zephyr replaces Linux. The two solve different problems. If your application needs a full UI stack, a Python runtime, containers or a rich networking and middleware ecosystem, Linux remains the natural choice. Where the priorities are deterministic real-time behaviour, fast boot, low power and a small, predictable footprint, Zephyr is purpose-built for the job. And its ecosystem is expanding quickly, steadily closing the gap in areas where Linux has traditionally led. It's a clean architecture, and it performs.
The Clock Speed Argument
Linux on PolarFire SoC runs the application cores at up to 600 MHz. That headroom is necessary, the OS overhead demands it to deliver acceptable responsiveness.
Zephyr does not carry that overhead. Running at 300 MHz, it delivers equal or better real-work throughput for the actual application logic, at half the dynamic power. Combined with the removal of DDR and its associated power delivery, the system-level efficiency improvement is significant.
For power-constrained or thermally-limited designs, this is not a secondary consideration. It is often a deciding one.
How the FPGA Fabric Is Your Secret Sauce
Removing DDR does not diminish what the FPGA side of PolarFire SoC can do. The fabric remains fully available for hardware acceleration, custom peripherals, pre-processing pipelines and, critically, for attaching alternative memory interfaces like HyperRAM when additional address space or frame buffer memory is needed.
This is the architecture that makes PolarFire SoC genuinely distinctive: a real multicore RISC-V processor, a capable FPGA fabric, and now—with Zephyr SMP—a software path that can exploit that combination without the overhead of a full operating system.
And remember that edge AI camera from earlier? On-device inference has real limits—especially CPU-only inference—in speed, in model size and in how much processor time it leaves for everything else. Say you needed more behind that model: faster inference or the same result without burning CPU cycles you'd rather spend elsewhere.
This is exactly where the fabric earns its name. Because it is an FPGA, you are not limited to bolting on more memory. You can drop a neural-network accelerator like CoreVectorBlox straight into the fabric, offload the CNN to dedicated hardware or add your own custom IP to extend the PolarFire SoC MSS however the application demands. If it fits and runs in the fabric, it is yours to build.
Is This the Right Architecture for Your Application?
Start with the bandwidth question. If your workload genuinely needs DDR-scale throughput, that answer is clear. But if the honest answer is "some external memory would help but we are nowhere near DDR saturation," or "we fit in scratchpad with careful design." then the design space opens up considerably.
This approach fits well where:
- Hard or soft real-time response is a primary requirement
- Power budget or thermal envelope is constrained
- BOM cost and PCB complexity are under pressure
- Deterministic, predictable multi-core execution matters
- The Linux ecosystem is not a hard requirement
Industrial control, motor drive, communications infrastructure, safety-critical edge nodes, these are the application spaces where a DDR-free or DDR-light design on PolarFire SoC with Zephyr SMP is worth evaluating seriously.
If your application needs the full Linux ecosystem—rich networking, complex file systems, broad userspace software support—the Linux-first path remains the right answer and is fully supported. The point is not that Zephyr replaces Linux . The point is that it is now a serious, first-class alternative for the applications where its strengths align, and that the memory architecture decision deserves to be made on actual requirements, not habit.
All images used courtesy of Microchip.