Packet Filtering with VPP on BlueField-3 Arm Cores and 100G NICs

FastNetMon

October 5, 2026

Home ‣ FastNetMon Blog ‣ Packet Filtering with VPP on BlueField-3 Arm Cores and 100G NICs
Portrait of a man with short dark hair and light stubble, wearing a dark gray T-shirt, facing the camera (circular crop).

This post is a repost of a technical blog originally published by Denys Haryachyy, shared here with permission as part of ongoing research and engineering work around FastNetMon’s inline traffic processing capabilities.

The NVIDIA BlueField-3 DPU carries sixteen Arm Cortex-A78AE cores right next to its ConnectX-7 network engine. I wanted to know how much an FD.io VPP data plane can filter when it runs on those Arm cores instead of the x86 host, so I built VPP with an ACL filter plugin for arm64, started it on the BlueField-3 in DPU mode and drove it with 2×100G of traffic. This post walks through every step of getting VPP up on the DPU — the card mode, the uplinks, hugepages, the container, the startup configuration — and then shows how fast it filters. In short: on a 100G port the Arm cores filter at line rate from 256-byte frames upward and on IMIX, while 64-byte floods top out at about 59 Mpps (42 Gbps, 41% of line rate) and 128-byte frames at 52 Mpps (64%).

Our Test Setup

The DPU sits in an AMD EPYC 9534 server. A separate Ryzen 9950X box runs TRex and is cabled to the two BlueField-3 uplinks with 100G DACs, no switch in between.

ComponentDetails
DPUNVIDIA BlueField-3, integrated ConnectX-7, 2×100G uplinks
Arm CPU16× Cortex-A78AE at 2.13 GHz, 512 KB L2 per core, 16 MB shared L3, 30 GB DDR5
DPU softwareBlueField bundle 2.9.1 (DOCA 24.11), Ubuntu 22.04, kernel 5.15.0-1057-bluefield
VPP25.10 built from source for arm64, DPDK mlx5 driver, in a Docker container on the Arm
GeneratorTRex 3.06 on a Ryzen 9 9950X with two ConnectX-5 Ex cards
VPP on the BlueField-3 Arm cores between the two uplinks
VPP owns both uplinks on the Arm side; management stays on oob_net0

Step 1 — Put the Card in DPU Mode

A BlueField-3 runs either as a plain NIC, where the x86 host owns the ports, or in DPU mode, where the embedded Arm system owns the uplinks and the host only sees representor-backed functions. The mode is the firmware parameter INTERNAL_CPU_MODEL: EMBEDDED_CPU(1) is DPU mode, SEPARATED_HOST(0) hands the ports to the host. The Arm OS ships mlxconfig, so I read and set the mode from the DPU itself, which keeps the host free of any Mellanox tools (NVIDIA: BlueField modes of operation).

sudo mlxconfig -d 03:00.0 -e q INTERNAL_CPU_MODEL
sudo mlxconfig -d 03:00.0 -y s INTERNAL_CPU_MODEL=1
sync

The query prints three columns — default, current and next boot — so it shows a pending change before it takes effect. The new value is applied only by a cold power cycle of the whole server: I power the chassis off through the BMC, wait for it to report off, power it on again and then check that the current column shows the new mode. A firmware reset or a warm reboot of the host leaves the old mode in place.

The DPU mode switch takes effect only after a cold power cycle of the host server, and that power cut also hits the Arm. Run sync on both sides first: files written to the Arm moments before the cut can come back empty.

Step 2 — Release the Uplinks From OVS

In DPU mode the Arm boots with Open vSwitch bridging each uplink to its host function: ovsbr1 joins p0 with pf0hpf, and ovsbr2 joins p1 with pf1hpf (NVIDIA: virtual switch on the DPU). DPDK needs the uplinks to itself, so I detach them from the bridges. The SSH session survives, because DPU management runs over the separate out-of-band port oob_net0, not over the uplinks.

sudo ovs-vsctl del-port ovsbr1 p0
sudo ovs-vsctl del-port ovsbr2 p1
# when finished, give the wire back to the host
sudo ovs-vsctl add-port ovsbr1 p0
sudo ovs-vsctl add-port ovsbr2 p1

Step 3 — Hugepages

The BlueField OS mounts only a 2 MB hugetlbfs at /dev/hugepages; there is no 1 GB mount, so the whole VPP setup uses 2 MB pages. 4,096 of them (8 GB) hold the 4 GB main heap and a one-million-buffer pool comfortably.

echo 4096 | sudo tee /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages

Step 4 — Run VPP in a Container

Nothing is compiled on the DPU. I build VPP 25.10 from the upstream source with make pkg-deb on a native arm64 GitHub Actions runner, inside an Ubuntu 24.04 container, and build the ACL filter plugin out of tree against those packages on the same runner. The resulting arm64 .deb files — vpp, libvppinfra, vpp-plugin-core, vpp-plugin-dpdk, vpp-crypto-engines and the plugin — are all the DPU needs.

The BlueField OS ships Docker, so VPP runs in a container on the Arm cores. I build the image on the DPU itself: Ubuntu 24.04 plus those packages, ibverbs-providers and libnuma1. The DPDK mlx5 driver talks to the NIC through rdma-core’s mlx5 provider and the /dev/infiniband verbs devices. mlx5 is a bifurcated driver, so the uplinks stay bound to the kernel mlx5_core driver — no vfio or igb_uio binding is needed (DPDK mlx5 guide: bifurcated driver).

docker run -d --name vpp --privileged --network host \
  -v /dev/hugepages:/dev/hugepages \
  -v /dev/infiniband:/dev/infiniband \
  -v /run/vpp:/run/vpp \
  -v $PWD/startup.conf:/etc/vpp/startup.conf \
  vpp-arm64 vpp -c /etc/vpp/startup.conf

The container needs --privileged and the host network so DPDK can open the PCI devices and the verbs interface, plus the hugepage mount and /run/vpp for the CLI and API sockets.

Alternative — FD.io Packages and an Out-of-Tree Plugin

VPP itself does not have to be compiled at all. FD.io publishes arm64 builds for Ubuntu 22.04 — the release the BlueField OS runs — and 24.04, and their DPDK plugin already contains the mlx5 driver. On top of them only the plugin is built, out of tree, against the vpp-dev headers and CMake helpers (FD.io 25.10 repository setup). The plugin is compiled against VPP’s internal headers, so I pin the exact package version and hold it, and rebuild the plugin whenever VPP moves.

curl -s https://packagecloud.io/install/repositories/fdio/2510/script.deb.sh | sudo bash
V=25.10.0-23~g744d3c715
sudo apt-get install -y vpp=$V vpp-plugin-core=$V vpp-plugin-dpdk=$V \
     vpp-dev=$V libvppinfra=$V libvppinfra-dev=$V \
     build-essential cmake ninja-build python3-ply
sudo apt-mark hold vpp vpp-plugin-core vpp-plugin-dpdk vpp-dev libvppinfra libvppinfra-dev

The plugin’s CMake project needs only two calls: find_package(VPP) loads the helpers that vpp-dev installs, and add_vpp_plugin generates the API code with vppapigen and compiles the data-plane node once per CPU variant (VPP sample plugin). vppapigen imports PLY, which is why python3-ply is in the list above.

cmake_minimum_required(VERSION 3.16)
project(acl-filter-plugin C)
find_package(VPP REQUIRED)
add_vpp_plugin(acl_filter
  SOURCES acl_filter.c acl_filter_api.c acl_filter_node.c
  MULTIARCH_SOURCES acl_filter_node.c
  API_FILES acl_filter.api
)

The packaged helpers install plugins under lib/vpp_plugins unless told otherwise, while the Debian VPP loads them from the multiarch directory. VPP_LIBRARY_DIR — not CMake’s usual CMAKE_INSTALL_LIBDIR — points the install at the right place:

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
      -DCMAKE_INSTALL_PREFIX=/usr -DVPP_LIBRARY_DIR=lib/$(gcc -print-multiarch)
cmake --build build
sudo cmake --install build      # /usr/lib/aarch64-linux-gnu/vpp_plugins/

I ran this end to end in an arm64 Ubuntu 22.04 container under QEMU on an x86 laptop: the packages installed, the plugin built and installed in about seven minutes of emulated time, VPP listed it with its own version string, and a drop rule matched and dropped all 1,000 UDP packets the packet generator sent through it. The image that results is a drop-in for the container in Step 4.

What this route gives up is CPU tuning. The packages are compiled for a generic -march=armv8-a+crc baseline plus a fixed set of node variants, and the plugin’s data-plane node inherits that same list from the packaged cpu.cmake (VPP cpu.cmake: aarch64 variants). None of them targets the Cortex-A78AE, so on the BlueField VPP runs the baseline code everywhere — the plugin’s node shows default as its only active variant. Flags such as -mcpu=cortex-a78ae, which DPDK recommends for the BlueField-3, cannot be applied here: VPP and its bundled DPDK arrive prebuilt, and adding them to the plugin alone would still leave the RX path, buffer handling and graph dispatch on the generic build. Tuning needs VPP and DPDK built from source.

FD.io packages + out-of-tree pluginVPP built from source
Build timeminutesabout 47 minutes per VPP version on a GitHub arm64 runner
Version controlpinned FD.io build, held in aptany tag, any patch set

Step 5 — The Startup Configuration

The uplinks are addressed by PCI address and renamed to p0 and p1 so the VPP interface names match the Linux ones. Each port gets as many RX queues as there are workers, and RSS hashes IPv4 and IPv6 so source-address variety spreads across all of them (VPP startup.conf: the dpdk section).

unix {
  cli-listen /run/vpp/cli.sock
  startup-config /etc/vpp/setup.vpp
  nodaemon
}
memory {
  main-heap-size 4G
  main-heap-page-size 2M
}
dpdk {
  no-multi-seg
  no-tx-checksum-offload
  dev 0000:03:00.0 { name p0 num-rx-queues 12 num-tx-queues 12 num-rx-desc 4096 num-tx-desc 4096 rss { ipv4 ipv6 } }
  dev 0000:03:00.1 { name p1 num-rx-queues 12 num-tx-queues 12 num-rx-desc 4096 num-tx-desc 4096 rss { ipv4 ipv6 } }
}
cpu {
  main-core 0
  corelist-workers 1-12
}
buffers { buffers-per-numa 1048576 }
plugins {
  plugin acl_filter_plugin.so { enable }
}

Twelve workers on cores 1–12 is the best split I measured: core 0 runs the VPP main thread, and cores 13–15 stay with Linux, OVS and sshd. Fourteen workers filtered less than twelve, because the extra workers compete for the same 16 MB L3 rather than adding throughput.

BlueField-3 Arm core layout for VPP
Core layout: main thread on core 0, twelve workers on cores 1–12, three cores left to the OS

Step 6 — Wire the Ports and the Filter

The startup script brings both uplinks up, cross-connects them at layer 2 in both directions and enables the ACL filter on each port, so every packet arriving on one uplink is classified and, unless a rule drops it, leaves on the other. An L2 cross-connect puts the ports in promiscuous mode, so the generator’s destination MAC does not matter. For a routed setup, address traffic to the uplink MACs of p0 and p1 — the host-side functions carry different MACs.

To measure, I read the NIC’s own rx_good_packets and rx_missed_errors counters from show hardware-interfaces at the start and end of a 15-second window. rx_missed_errors rising means the Arm cores are saturated and the good-packet rate is the ceiling; when it stays at zero the DPU keeps up with the wire.

How Fast It Filters

Each point below is a 15-second sample with a drop rule matching the flood, one 100G port in and twelve workers, against the 100G line rate for that frame size.

Drop rate per frame size on the BlueField-3 Arm cores
Drop rate per frame size, one 100G ingress port
FrameDrop, 1 portForward, 1 portDrop, 2 portsForward, 2 ports
64 B60 Mpps · 42 Gbps50 Mpps · 35 Gbps59 Mpps · 42 Gbps66 Mpps · 46 Gbps
IMIX33.8 Mpps · ~100 Gbps (line rate)33.8 Mpps · ~100 Gbps (line rate)44 Mpps · 133 Gbps39 Mpps · 118 Gbps
1500 B8.1 Mpps · ~100 Gbps (line rate)8.1 Mpps · ~100 Gbps (line rate)16 Mpps · ~200 Gbps (line rate)15 Mpps · 183 Gbps

The 64-byte limit is shared by all twelve workers: flooding both uplinks reaches the same ~60 Mpps in aggregate, while large frames scale to both ports at line rate.

Two Ports: Pin RX Queues per Port

With both uplinks loaded, VPP’s default placement spreads every port’s queues across every worker, so each worker polls both ports and the two flows of buffers thrash the same caches. For two-port runs I configure six RX queues per port and pin them with set interface rx-placement: queues of p0 to the first six workers, queues of p1 to the other six (VPP CLI: set interface rx-placement). That restores the full 59 Mpps aggregate.

64-byte drop rate with one port, two ports default placement and two ports dedicated placement
64-byte drop rate: dedicated per-port placement recovers the full rate

Working Set: Flows Matter, Rule Count Does Not

The table can be large: I loaded 983,045 rules and the per-packet cost stayed the same as with 10,000. What moves the rate is how many different rules the traffic touches at once. The BlueField-3 holds its peak up to about a hundred simultaneous flows and then declines steadily as the working set outgrows the 512 KB L2 and 16 MB L3.

64-byte drop rate versus simultaneous active flows
64-byte drop rate versus simultaneous active flows, 983,045 rules installed
ChangePer-packet costEffect
10,000 → 983,000 rules (one match shape)29.1 → 29.1 ticksnone — hash lookup, untouched rules never reach the cache
+5 port-range rules (2 match shapes)29.1 → 47.7 ticks+64% — one more hash table probed per packet
1,000 → 10,000 active flows29.1 → 56.6 ticks×2 — the working set leaves the cache

Reading Cycle Counters on Arm

The clock column of show runtime counts ticks of the Arm generic timer, which runs at 330 MHz on the BlueField-3, not CPU cycles. The A78AE cores run at 2.13 GHz (measured with perf stat), so one VPP tick is about 6.5 core cycles. The 29 ticks per packet above are therefore roughly 190 core cycles, and only the converted figure compares with the TSC-based clocks VPP reports on x86.

Summary

VPP runs on the BlueField-3 Arm cores in a Docker container, from arm64 packages built natively in CI. The card goes into DPU mode through mlxconfig on the Arm and a cold power cycle; the uplinks come out of the OVS bridges while management stays on oob_net0; VPP uses 2 MB hugepages and the bifurcated mlx5 driver from a privileged container; twelve workers with one RSS queue each give the best rate. On that setup the DPU filters at 100G line rate for frames of 256 bytes and larger and for IMIX; smaller frames are CPU-bound, at about 59 Mpps (42 Gbps) for 64 bytes and 52 Mpps (64 Gbps) for 128 bytes. With both uplinks loaded, 1500-byte frames reach about 200 Gbps. The filter keeps its per-packet cost flat from ten thousand to nearly a million rules.

References