Edge AI is the practice of running machine learning inference where the sensor and the user are: on microcontrollers, mobile SoCs, cameras and gateways rather than in a datacenter. The discipline spans TinyML, model quantization, model pruning, low-latency local inference and the runtime stack that makes device-side model execution possible, from TensorRT to ONNX Runtime to vendor NPU SDKs. MLPerf Tiny v1.4 drew 25 system configurations from nine organizations and measured accuracy, latency and energy on workloads whose reference models run under 100 kB . MLCommons added a streaming wake-word benchmark for continuous audio, the shape most real edge deployments take .
Challenges in Edge AI Recruiting
Low-latency local inference splits work by latency budget, not model size
The phrase edge covers tiers with nothing in common. A Qualcomm submission on the Snapdragon 8 Elite Gen 5 Sensing Hub reported keyword spotting, visual wake words and image classification under 0.30 ms per inference, a regime where the model barely registers against the scheduler . An automotive perception stack on a Jetson-class SoC budgets tens of milliseconds and carries a GPU memory planner. An industrial gateway runs the same model family at whatever latency its network allows. These are different disciplines: interrupt handlers versus tensor memory versus throughput tuning. The v1.4 round surfaced the spread directly, from a 22.2-microjoule image classification on a dedicated NPU to a Versal platform pushing over 1,800 inferences per second through a custom on-fabric engine . The CV that says deployed models to edge devices does not say which tier the candidate worked, and employers frequently interview one when they need another.
TinyML runs on microcontroller flash, RAM and joules
Below the SoC tier sits TinyML, where the reference models of the MLPerf Tiny suite stay under 100 kB because the target is a Cortex-M-class microcontroller, not a GPU . The suite grew from four tasks in 2021 to five in v1.3, adding a streaming wake-word task that scores false accepts and false rejects over continuous audio rather than a single clip . The v1.4 round made energy a first-class measurement: every submission must hit a fixed quality target, then competes on latency and energy per inference measured through calibrated power harnesses . Syntiant's NDP120 ran the streaming wake-word benchmark at a 3.3% duty cycle, idle nearly 97% of the time with capacity left for noise cancellation . Server-side ML engineers arrive at these constraints like tourists. Flash layout, fixed-point kernels, memory arena planning and sleep states are the job, and none of them appears in a PyTorch tutorial. The candidate pool instead comes from embedded software, silicon vendors and a small band of researchers who treat energy budgets as the primary optimization target.
Model quantization is a per-accelerator craft
Model quantization is where the portability promise of edge AI ends. TensorRT supports INT8, FP8, INT4 and FP4 through post-training and quantization-aware workflows, with Q/DQ nodes marking exactly where conversions happen in the graph . In the INT8 world that still dominates deployed vision models, the work is calibration: choosing the representative set, running the activation ranges, checking which layers degrade and by how much. The same model then meets hardware with different rules, which is why the craft does not transfer between targets. A quantization engineer who has tuned one accelerator has learned one flavor of a family: the calibration data, the per-layer sensitivity, the fallback graph. Treating that experience as portable is how deployments ship a model that is smaller and slower.
Mobile neural engines pin teams to NPU schedulers and vendor SDKs
ONNX Runtime reaches the NPUs through execution providers: NNAPI on Android, CoreML on Apple platforms, the Qualcomm QNN provider for Hexagon, XNNPACK for Arm CPUs, and the TensorRT provider on Jetson-class GPUs . Each provider is a different submission to a different scheduler. Qualcomm's HTP backend only accepts quantized graphs, so a float model must go through a quantization pass before the NPU will touch it . The vendor's model library lists which precision each chipset supports, from FP16 on newer Hexagon silicon to INT8 and INT16 across the board . Engineers who have shipped here know their accelerators the way compiler people know their targets: fallback behavior, memory regions, vendor toolchains, and the profiling tools that show where a graph actually runs, since a session can silently route work to the CPU while the logs keep saying NPU. The title mobile ML engineer hides which of these stacks a candidate owns.
ONNX portability ends where execution providers begin
ONNX is an interchange format, not a runtime contract. The graph exports cleanly, then each execution provider partitions it differently: TensorRT fuses layers and needs shape information to build an engine, NNAPI falls back to its CPU reference for unsupported operators, and QNN drops to CPU for anything outside its op set . Provider priority is itself a design decision, since the same session can prefer one backend and degrade to another per node. A model that runs everywhere in PyTorch becomes a per-target integration project the day it leaves the notebook. Candidates who have done edge deployment own those fallback decisions: which operators stayed on the NPU, what moved to CPU, and what that did to the latency budget. Candidates who have only exported ONNX once cannot answer the question.
Device-side model execution claims die on the latency-power budget
The verification questions for device-side model execution are the same ones the benchmarks enforce: which accelerator, which precision, which latency, at which power. Ask what the model was quantized with and what the calibration set contained . Ask for the measured inference time on the shipped silicon, not the reference board. Ask what happens when the engine cannot be built: does the candidate know what TensorRT engine caching does, or which QNN operators failed over to CPU . A measured delta is the standard: a TensorRT FP8 engine on an RTX 6000 Ada cut the CLIP image encoder from 166.2 ms to 119.8 ms with the checkpoint exported through Q/DQ nodes, and a deployment engineer should be able to produce the equivalent numbers for their own stack . A solid answer names numbers and fallbacks; a paper answer names frameworks. A model pruning story gets the same treatment: which sparsity, which structure, and what the accuracy-versus-footprint curve looked like after each round, rather than a percentage quoted from a paper. The cost of a miss is a device program that misses its launch because the model draws twice the allowed power or wakes the SoC for ten milliseconds it does not have. That failure is found at the last integration gate, where rework costs a release rather than a sprint.
References
- MLPerf Tiny: Benchmarking AI at the Edge (v1.4 Results) — MLCommons. (accessed 2026-09-28)
- MLCommons New MLPerf Tiny 1.3 Benchmark Results Released — MLCommons. (accessed 2026-09-28)
- Quantization Workflows — NVIDIA TensorRT Documentation. (accessed 2026-09-28)
- ONNX Runtime Execution Providers — ONNX Runtime Documentation. (accessed 2026-09-28)
- QNN Execution Provider — ONNX Runtime Documentation (GitHub). (accessed 2026-09-28)
- Qualcomm AI Hub Models: Model Support Data — Qualcomm (GitHub). (accessed 2026-09-28)
- Model Quantization: Turn FP8 Checkpoints into High-Performance Inference Engines with NVIDIA TensorRT — NVIDIA Technical Blog. (accessed 2026-09-28)
