A developer describes building a custom FPGA-based TPU for matrix operations, progressing from a Zynq board to a smaller KC705 with VexRiscv CPU and 4×4 int8 cores. The system compiles JAX programs to C, running compute-intensive operations on the arrays while keeping loops on the CPU, achieving 104 ms inference on MNIST classification.