
This post is a repost of a technical blog originally published by Denys Haryachyy, shared here with permission as part of ongoing research and engineering work around FastNetMon’s inline traffic processing capabilities.
The NVIDIA BlueField-3 DPU carries sixteen Arm Cortex-A78AE cores right next to its ConnectX-7 network engine. I wanted to know how much an FD.io VPP data plane can filter when it runs on those Arm cores instead of the x86 host, so I built VPP with an ACL filter plugin for arm64, started it on the BlueField-3 in DPU mode and drove it with 2×100G of traffic. This post walks through every step of getting VPP up on the DPU — the card mode, the uplinks, hugepages, the container, the startup configuration — and then shows how fast it filters. In short: on a 100G port the Arm cores filter at line rate from 256-byte frames upward and on IMIX, while 64-byte floods top out at about 59 Mpps (42 Gbps, 41% of line rate) and 128-byte frames at 52 Mpps (64%).
Our Test Setup
The DPU sits in an AMD EPYC 9534 server. A separate Ryzen 9950X box runs TRex and is cabled to the two BlueField-3 uplinks with 100G DACs, no switch in between.
| Component | Details |
|---|---|
| DPU | NVIDIA BlueField-3, integrated ConnectX-7, 2×100G uplinks |
| Arm CPU | 16× Cortex-A78AE at 2.13 GHz, 512 KB L2 per core, 16 MB shared L3, 30 GB DDR5 |
| DPU software | BlueField bundle 2.9.1 (DOCA 24.11), Ubuntu 22.04, kernel 5.15.0-1057-bluefield |
| VPP | 25.10 built from source for arm64, DPDK mlx5 driver, in a Docker container on the Arm |
| Generator | TRex 3.06 on a Ryzen 9 9950X with two ConnectX-5 Ex cards |

Step 1 — Put the Card in DPU Mode
A BlueField-3 runs either as a plain NIC, where the x86 host owns the ports, or in DPU mode, where the embedded Arm system owns the uplinks and the host only sees representor-backed functions. The mode is the firmware parameter INTERNAL_CPU_MODEL: EMBEDDED_CPU(1) is DPU mode, SEPARATED_HOST(0) hands the ports to the host. The Arm OS ships mlxconfig, so I read and set the mode from the DPU itself, which keeps the host free of any Mellanox tools (NVIDIA: BlueField modes of operation).
sudo mlxconfig -d 03:00.0 -e q INTERNAL_CPU_MODEL sudo mlxconfig -d 03:00.0 -y s INTERNAL_CPU_MODEL=1 sync
The query prints three columns — default, current and next boot — so it shows a pending change before it takes effect. The new value is applied only by a cold power cycle of the whole server: I power the chassis off through the BMC, wait for it to report off, power it on again and then check that the current column shows the new mode. A firmware reset or a warm reboot of the host leaves the old mode in place.
The DPU mode switch takes effect only after a cold power cycle of the host server, and that power cut also hits the Arm. Run sync on both sides first: files written to the Arm moments before the cut can come back empty.
Step 2 — Release the Uplinks From OVS
In DPU mode the Arm boots with Open vSwitch bridging each uplink to its host function: ovsbr1 joins p0 with pf0hpf, and ovsbr2 joins p1 with pf1hpf (NVIDIA: virtual switch on the DPU). DPDK needs the uplinks to itself, so I detach them from the bridges. The SSH session survives, because DPU management runs over the separate out-of-band port oob_net0, not over the uplinks.
sudo ovs-vsctl del-port ovsbr1 p0 sudo ovs-vsctl del-port ovsbr2 p1 # when finished, give the wire back to the host sudo ovs-vsctl add-port ovsbr1 p0 sudo ovs-vsctl add-port ovsbr2 p1
Step 3 — Hugepages
The BlueField OS mounts only a 2 MB hugetlbfs at /dev/hugepages; there is no 1 GB mount, so the whole VPP setup uses 2 MB pages. 4,096 of them (8 GB) hold the 4 GB main heap and a one-million-buffer pool comfortably.
echo 4096 | sudo tee /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages
Step 4 — Run VPP in a Container
Nothing is compiled on the DPU. I build VPP 25.10 from the upstream source with make pkg-deb on a native arm64 GitHub Actions runner, inside an Ubuntu 24.04 container, and build the ACL filter plugin out of tree against those packages on the same runner. The resulting arm64 .deb files — vpp, libvppinfra, vpp-plugin-core, vpp-plugin-dpdk, vpp-crypto-engines and the plugin — are all the DPU needs.
The BlueField OS ships Docker, so VPP runs in a container on the Arm cores. I build the image on the DPU itself: Ubuntu 24.04 plus those packages, ibverbs-providers and libnuma1. The DPDK mlx5 driver talks to the NIC through rdma-core’s mlx5 provider and the /dev/infiniband verbs devices. mlx5 is a bifurcated driver, so the uplinks stay bound to the kernel mlx5_core driver — no vfio or igb_uio binding is needed (DPDK mlx5 guide: bifurcated driver).
docker run -d --name vpp --privileged --network host \ -v /dev/hugepages:/dev/hugepages \ -v /dev/infiniband:/dev/infiniband \ -v /run/vpp:/run/vpp \ -v $PWD/startup.conf:/etc/vpp/startup.conf \ vpp-arm64 vpp -c /etc/vpp/startup.conf
The container needs --privileged and the host network so DPDK can open the PCI devices and the verbs interface, plus the hugepage mount and /run/vpp for the CLI and API sockets.
Alternative — FD.io Packages and an Out-of-Tree Plugin
VPP itself does not have to be compiled at all. FD.io publishes arm64 builds for Ubuntu 22.04 — the release the BlueField OS runs — and 24.04, and their DPDK plugin already contains the mlx5 driver. On top of them only the plugin is built, out of tree, against the vpp-dev headers and CMake helpers (FD.io 25.10 repository setup). The plugin is compiled against VPP’s internal headers, so I pin the exact package version and hold it, and rebuild the plugin whenever VPP moves.
curl -s https://packagecloud.io/install/repositories/fdio/2510/script.deb.sh | sudo bash
V=25.10.0-23~g744d3c715
sudo apt-get install -y vpp=$V vpp-plugin-core=$V vpp-plugin-dpdk=$V \
vpp-dev=$V libvppinfra=$V libvppinfra-dev=$V \
build-essential cmake ninja-build python3-ply
sudo apt-mark hold vpp vpp-plugin-core vpp-plugin-dpdk vpp-dev libvppinfra libvppinfra-dev
The plugin’s CMake project needs only two calls: find_package(VPP) loads the helpers that vpp-dev installs, and add_vpp_plugin generates the API code with vppapigen and compiles the data-plane node once per CPU variant (VPP sample plugin). vppapigen imports PLY, which is why python3-ply is in the list above.
cmake_minimum_required(VERSION 3.16) project(acl-filter-plugin C) find_package(VPP REQUIRED) add_vpp_plugin(acl_filter SOURCES acl_filter.c acl_filter_api.c acl_filter_node.c MULTIARCH_SOURCES acl_filter_node.c API_FILES acl_filter.api )
The packaged helpers install plugins under lib/vpp_plugins unless told otherwise, while the Debian VPP loads them from the multiarch directory. VPP_LIBRARY_DIR — not CMake’s usual CMAKE_INSTALL_LIBDIR — points the install at the right place:
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_INSTALL_PREFIX=/usr -DVPP_LIBRARY_DIR=lib/$(gcc -print-multiarch)
cmake --build build
sudo cmake --install build # /usr/lib/aarch64-linux-gnu/vpp_plugins/
I ran this end to end in an arm64 Ubuntu 22.04 container under QEMU on an x86 laptop: the packages installed, the plugin built and installed in about seven minutes of emulated time, VPP listed it with its own version string, and a drop rule matched and dropped all 1,000 UDP packets the packet generator sent through it. The image that results is a drop-in for the container in Step 4.
What this route gives up is CPU tuning. The packages are compiled for a generic -march=armv8-a+crc baseline plus a fixed set of node variants, and the plugin’s data-plane node inherits that same list from the packaged cpu.cmake (VPP cpu.cmake: aarch64 variants). None of them targets the Cortex-A78AE, so on the BlueField VPP runs the baseline code everywhere — the plugin’s node shows default as its only active variant. Flags such as -mcpu=cortex-a78ae, which DPDK recommends for the BlueField-3, cannot be applied here: VPP and its bundled DPDK arrive prebuilt, and adding them to the plugin alone would still leave the RX path, buffer handling and graph dispatch on the generic build. Tuning needs VPP and DPDK built from source.
| FD.io packages + out-of-tree plugin | VPP built from source | |
|---|---|---|
| Build time | minutes | about 47 minutes per VPP version on a GitHub arm64 runner |
| Version control | pinned FD.io build, held in apt | any tag, any patch set |
Step 5 — The Startup Configuration
The uplinks are addressed by PCI address and renamed to p0 and p1 so the VPP interface names match the Linux ones. Each port gets as many RX queues as there are workers, and RSS hashes IPv4 and IPv6 so source-address variety spreads across all of them (VPP startup.conf: the dpdk section).
unix {
cli-listen /run/vpp/cli.sock
startup-config /etc/vpp/setup.vpp
nodaemon
}
memory {
main-heap-size 4G
main-heap-page-size 2M
}
dpdk {
no-multi-seg
no-tx-checksum-offload
dev 0000:03:00.0 { name p0 num-rx-queues 12 num-tx-queues 12 num-rx-desc 4096 num-tx-desc 4096 rss { ipv4 ipv6 } }
dev 0000:03:00.1 { name p1 num-rx-queues 12 num-tx-queues 12 num-rx-desc 4096 num-tx-desc 4096 rss { ipv4 ipv6 } }
}
cpu {
main-core 0
corelist-workers 1-12
}
buffers { buffers-per-numa 1048576 }
plugins {
plugin acl_filter_plugin.so { enable }
}
Twelve workers on cores 1–12 is the best split I measured: core 0 runs the VPP main thread, and cores 13–15 stay with Linux, OVS and sshd. Fourteen workers filtered less than twelve, because the extra workers compete for the same 16 MB L3 rather than adding throughput.

Step 6 — Wire the Ports and the Filter
The startup script brings both uplinks up, cross-connects them at layer 2 in both directions and enables the ACL filter on each port, so every packet arriving on one uplink is classified and, unless a rule drops it, leaves on the other. An L2 cross-connect puts the ports in promiscuous mode, so the generator’s destination MAC does not matter. For a routed setup, address traffic to the uplink MACs of p0 and p1 — the host-side functions carry different MACs.
To measure, I read the NIC’s own rx_good_packets and rx_missed_errors counters from show hardware-interfaces at the start and end of a 15-second window. rx_missed_errors rising means the Arm cores are saturated and the good-packet rate is the ceiling; when it stays at zero the DPU keeps up with the wire.
How Fast It Filters
Each point below is a 15-second sample with a drop rule matching the flood, one 100G port in and twelve workers, against the 100G line rate for that frame size.

| Frame | Drop, 1 port | Forward, 1 port | Drop, 2 ports | Forward, 2 ports |
|---|---|---|---|---|
| 64 B | 60 Mpps · 42 Gbps | 50 Mpps · 35 Gbps | 59 Mpps · 42 Gbps | 66 Mpps · 46 Gbps |
| IMIX | 33.8 Mpps · ~100 Gbps (line rate) | 33.8 Mpps · ~100 Gbps (line rate) | 44 Mpps · 133 Gbps | 39 Mpps · 118 Gbps |
| 1500 B | 8.1 Mpps · ~100 Gbps (line rate) | 8.1 Mpps · ~100 Gbps (line rate) | 16 Mpps · ~200 Gbps (line rate) | 15 Mpps · 183 Gbps |
The 64-byte limit is shared by all twelve workers: flooding both uplinks reaches the same ~60 Mpps in aggregate, while large frames scale to both ports at line rate.
Two Ports: Pin RX Queues per Port
With both uplinks loaded, VPP’s default placement spreads every port’s queues across every worker, so each worker polls both ports and the two flows of buffers thrash the same caches. For two-port runs I configure six RX queues per port and pin them with set interface rx-placement: queues of p0 to the first six workers, queues of p1 to the other six (VPP CLI: set interface rx-placement). That restores the full 59 Mpps aggregate.

Working Set: Flows Matter, Rule Count Does Not
The table can be large: I loaded 983,045 rules and the per-packet cost stayed the same as with 10,000. What moves the rate is how many different rules the traffic touches at once. The BlueField-3 holds its peak up to about a hundred simultaneous flows and then declines steadily as the working set outgrows the 512 KB L2 and 16 MB L3.

| Change | Per-packet cost | Effect |
|---|---|---|
| 10,000 → 983,000 rules (one match shape) | 29.1 → 29.1 ticks | none — hash lookup, untouched rules never reach the cache |
| +5 port-range rules (2 match shapes) | 29.1 → 47.7 ticks | +64% — one more hash table probed per packet |
| 1,000 → 10,000 active flows | 29.1 → 56.6 ticks | ×2 — the working set leaves the cache |
Reading Cycle Counters on Arm
The clock column of show runtime counts ticks of the Arm generic timer, which runs at 330 MHz on the BlueField-3, not CPU cycles. The A78AE cores run at 2.13 GHz (measured with perf stat), so one VPP tick is about 6.5 core cycles. The 29 ticks per packet above are therefore roughly 190 core cycles, and only the converted figure compares with the TSC-based clocks VPP reports on x86.
Summary
VPP runs on the BlueField-3 Arm cores in a Docker container, from arm64 packages built natively in CI. The card goes into DPU mode through mlxconfig on the Arm and a cold power cycle; the uplinks come out of the OVS bridges while management stays on oob_net0; VPP uses 2 MB hugepages and the bifurcated mlx5 driver from a privileged container; twelve workers with one RSS queue each give the best rate. On that setup the DPU filters at 100G line rate for frames of 256 bytes and larger and for IMIX; smaller frames are CPU-bound, at about 59 Mpps (42 Gbps) for 64 bytes and 52 Mpps (64 Gbps) for 128 bytes. With both uplinks loaded, 1500-byte frames reach about 200 Gbps. The filter keeps its per-packet cost flat from ten thousand to nearly a million rules.
References
- NVIDIA DOCA: BlueField modes of operation and INTERNAL_CPU_MODEL
- NVIDIA BlueField BSP: virtual switch (OVS) on the DPU
- DPDK: NVIDIA BlueField board support package guide
- DPDK mlx5 driver: bifurcated driver model
- NVIDIA NICs performance report with DPDK 25.03 (BlueField-3 Arm results)
- VPP 25.10 startup.conf reference: the dpdk section
- VPP 25.10 CLI: set interface rx-placement
- FD.io VPP 25.10 packages, including arm64
- VPP sample plugin: out-of-tree plugin layout
- VPP cpu.cmake: aarch64 multiarch variants and baseline flags
- VPP docs: multiarch node function variants
- FD.io wiki: VPP on AArch64
- Kashyap et al., Understanding the Idiosyncrasies of Emerging BlueField DPUs (ICS 2025)
- Schrötter et al., XenoFlow: How Fast Can a SmartNIC-Based DNS Load Balancer Run? (BlueField-3)
- BlueField DPU setup notes for non-DGX servers (j3soon)
- Red Hat: offloading network functions to BlueField-2 DPUs
- Arm Dataplane Stack: VPP reference solutions on Arm Neoverse






