Per-pod visibility into NPU and accelerator usage on Kubernetes.

KubeNPU counts accelerator calls per pod with eBPF and exports them to Prometheus.

Every DRM and accel driver routes userspace through one kernel function,

drm_ioctl. The agent attaches a fentry program there and filters in-kernel:

ffmpeg in a pod

└─ ioctl(fd, DRM_IOCTL_I915_GEM_EXECBUFFER2)

└─ drm_ioctl()

└─ fentry program

├─ filp->f_inode->i_rdev → 226:128

├─ devices[226:128] → vendor id, or drop

├─ cmd & 0xff → 0x69

├─ ioctls[vendor, 0x69] → kind=submit, or drop

└─ ringbuf ← {timestamp, cgroup_id, tgid, cmd, major, minor, kind}

Both maps are filled by Go before the program is attached: device discovery

walks /dev/dri and sysfs, the vendor is matched by driver name, and its ioctl

table is uploaded. More than half the calls never reach userspace.

Userspace then turns cgroup_id into a pod name:

cgroup_id 117316

→ /sys/fs/cgroup/kubepods.slice/.../cri-containerd-079adc….scope

→ CRI ContainerStatus

→ default / ffmpeg-vaapi-test / ffmpeg

The cgroup index is rebuilt on a miss, not on a timer a pod created after startup is picked up within one event.

This works identically across vendors because ioctl numbers are frozen uAPI. Unlike kernel function names, they do not change between releases.

kubenpu_ioctl_total{pod,namespace,container,device,vendor,kind="submit|alloc|wait"}

kubenpu_device_info{device,vendor,driver,pci_id,numa_node}

kubenpu_events_dropped_total{reason}

kubenpu_cgroup_index_rebuilds_total

device is the PCI address, which survives reboots. card1 does not, because

its minor number can change. The node name is in kubenpu_device_info.

Counts are exact, never estimated. They are call counts, not utilization: on Intel UHD 620 a workload with 4× more submissions used 7× less device time.

Counting starts with the agent, not with the pod. Anything a pod did before

that is missing, mostly kind="alloc", which happens once at startup. The same

applies after every DaemonSet rollout.

docker build -t kubenpu-local .

docker run --rm --name kubenpu --privileged \

-p 127.0.0.1:8080:8080 \

-v /sys/fs/cgroup:/sys/fs/cgroup:ro \

-v /sys:/sys:ro -v /dev:/dev:ro \

-v /run/k3s/containerd:/run/k3s/containerd:ro \

kubenpu-localcurl -s localhost:8080/metrics | grep kubenpu_--privileged is only for the quick test. In Kubernetes the agent runs with

CAP_BPF and CAP_PERFMON.

make build

sudo ./bin/agent -debug-debug turns on the full verifier log - without it a rejected program prints

two useless lines.

To check what the agent sees before touching the kernel:

make build-kubenpuctl

./bin/kubenpuctl devicesDRIVER PCI ID ADDRESS NODES VENDOR

i915 8086:5917 0000:00:02.0 226:1,226:128 i915

This reads sysfs only and needs no privileges. If your device shows up with

- in the VENDOR column, KubeNPU found the hardware but has no implementation

for it yet.

bpf/vmlinux.h is checked in and works on any kernel with BTF, since the

program is CO-RE and relocates field offsets at load time. To regenerate it

from your own kernel:

bpftool btf dump file /sys/kernel/btf/vmlinux format c > bpf/vmlinux.h

make generate

make buildhelm install kubenpu deploy/helm -n kubenpu --create-namespacekubectl -n kubenpu port-forward daemonset/kubenpu 8080:8080

curl -s localhost:8080/metrics | grep kubenpu_ioctl_totalSet serviceMonitor.enabled=true if you run the Prometheus Operator.

kubectl apply -k deploy/kustomize/baseOptional components: deploy/kustomize/components/servicemonitor and

deploy/kustomize/components/accel.

cd deploy/tanka

tk env set environments/default --server=<api-server-url>

tk apply environments/defaultOverride _config in environments/default/main.jsonnet, same keys as Helm values.yaml.

- Linux 6.1+ with BTF (CONFIG_DEBUG_INFO_BTF=y)

- cgroup v2, with either the systemd or the cgroupfs cgroup driver

- a CRI v1 runtime, socket reachable by the agent

Check that the node uses cgroup v2:

ls /sys/fs/cgroup/cgroup.controllersBoth kubelet cgroup drivers are supported: kubepods.slice under

/sys/fs/cgroup is the systemd driver, plain kubepods is cgroupfs.

CUDA does not go through drm_ioctl, so nothing KubeNPU does applies to it.

Use DCGM exporter.

Copy pkg/hw/ivpu/ and fill in the numbers from the kernel's uapi header.

- More hardware, starting with accelerators used in production, such as

Qualcomm Cloud AI 100 (qaic, available on AWS DL2q) and Intel Gaudi (habanalabs). Both drivers go throughdrm_ioctl.

- TPUs and other accelerators whose drivers do not use drm_ioctl. This needs a different hook and is still being researched.

- sched_extplacement hints: run each task on CPUs close to the accelerator it uses, without changing the application.

- BPF LSM, so KubeNPU could also allow or block a pod's access to a device, not only observe it.

Apache 2.0. The eBPF program under bpf/ is GPL the helpers it uses are

GPL-only, and the program will not load otherwise.