Computer vision is the discipline of extracting structure from pixels: the detectors, segmenters, trackers, and inspectors that let machines read images and video the way they read text. Its crafts split by what the pixel feeds: object detection and image recognition for what is present, image segmentation for where it is, 3D vision for how it sits in space, visual inspection and video analysis for how it behaves over time. The field got its foundation model in April 2023, when Meta released the Segment Anything Model trained on over a billion masks across 11 million images, with zero-shot transfer competitive with fully supervised systems . That shift, from training a model per task to prompting one model for many tasks, now defines the skill market.
Challenges in Computer Vision Recruiting
Image segmentation got its foundation model
SAM changed what segmentation work is. Released in April 2023, the Segment Anything Model was trained on over a billion masks across 11 million licensed images and transfers zero-shot to tasks like edge detection and instance segmentation with quality often competitive with fully supervised baselines . The craft moved accordingly. Five years ago a segmentation engineer trained a U-Net on a custom annotation set; today the same engineer prompts a pretrained model and spends the real effort on mask quality control, edge fidelity, and the domain gaps the foundation model never saw. That is the hiring split. A candidate who has only built supervised segmentation pipelines holds half the modern stack; a candidate who has only prompted SAM cannot diagnose a bad mask. The scarce profile does both, and most segmentation seats quietly need exactly that profile. The dataset side matters too: SAM's masks came from a model-in-the-loop data engine that grew annotation into a machine learning problem of its own, and candidates who have run annotation loops understand quality costs that pure model builders never see .
Object detection runs on a twenty-year lineage
Object detection is old enough to have its own canon, and hiring ignores it at a cost. NeurIPS 2025 gave its Test of Time award to Faster R-CNN, the 2015 paper that replaced handcrafted proposals with a region proposal network and has been cited more than 56,700 times . The lineage matters because detectors still rest on that furniture: proposals, anchors or their modern equivalents, non-maximum suppression, and mAP computed at fixed IoU thresholds. Candidates who learned detection through YOLO wrappers and pretrained checkpoints often cannot explain why their detector double-counts overlapping objects, or what happens to recall when the IoU threshold moves. The engineer who can is the one you want when the product's failure mode is a missed box rather than a wrong label, which in production it almost always is.
Visual inspection is a lighting problem with a model attached
Machine vision on the factory floor is still a systems craft, and the market is growing: A3's tracking shows North American machine vision transactions up 8.8% year over year in Q2 2026, with systems up 9.5% and components up 4.6% . Visual inspection hires, though, are made or broken upstream of the model. Part presentation, lighting geometry, lens choice, vibration, and line speed decide whether a defect is even visible in the frame; the classifier only decides what happens next. A candidate who has tuned a detection model on public data and one who has brought an inspection cell up on a moving line share a resume keyword and nothing else. The interview that matters asks about one production inspection failure: what changed, optics or model, and how they proved which.
Image recognition turned into open-vocabulary matching
Image recognition no longer ends in a fixed set of classes. Open-vocabulary models trained on image-text pairs match pixels to concepts the training set never enumerated, so the recognition engineer's artifact is now the embedding space and the prompts that query it, not the softmax head. That redefines the role. The old skill set, architecture design and class balancing, still matters for low-resource and edge cases; the new skill set, concept curation and alignment, decides whether the system finds the thing the customer described in their own words. Candidates split between the two populations, and products that need both usually hire one and discover the gap after launch, when a reasonable-sounding concept string returns nothing and nobody can explain why. The verification probe follows the split: ask how the candidate would retrieve an object the training vocabulary cannot name, and whether the answer involves a new model, a new embedding, or a new prompt. Only the middle option is worth hiring for in a product that will see inputs the team never anticipated.
Video analysis added temporal memory to the stack
Video analysis stopped being image analysis repeated per frame. SAM 2 extended the segmentation model to video with streaming inference and a memory module, tracking objects across frames, including through temporary disappearance, and doing it fast enough for real-time use . The temporal memory is the craft change: identity now persists across time, which is what tracking, action understanding, and long-form analysis all demand. Hiring for video analysis therefore needs evidence of temporal reasoning, not just per-frame accuracy. A candidate who has built video systems can discuss ID swaps, memory budgets, and what breaks at the cut; a candidate whose video experience is running a still-image model on frames cannot. That difference is the entire distance between the two job titles.
3D vision split between reconstruction and rendering
3D vision now has two economies. NeRF-style radiance fields optimize a scene into a neural network; 3D Gaussian splatting, from 2023, instead represents scenes as millions of anisotropic Gaussians with a tile-based differentiable rasterizer, reaching real-time rendering at 1080p with quality matching the previous best . The hiring consequence is a genuine fork. Reconstruction engineers think in calibration, multi-view geometry, and point clouds; rendering engineers think in splats, visibility, and frame budgets. Robotic perception, mapping, and inspection consume the first population; capture, synthesis, and graphics consume the second. A job description that says 3D vision without naming which economy the product runs in will get half its shortlist from the wrong side. The screening question that fixes it is one sentence: what was the last 3D artifact the candidate shipped, a reconstructed point cloud, a rendered frame, or a depth map, and which one the product actually consumes downstream.
Vision-based AI claims fail under the camera question
Verification in this discipline starts at the lens. A candidate claiming vision-based AI depth should be able to walk through a real system from acquisition to decision: the sensor, the optics, the lighting, the calibration, the model, and the metric, plus one production failure and what fixed it. The question that separates populations is embarrassingly simple: how did the images get into your model, and what could change between the lab image and the production image? Model-centric candidates hand back a story about the training set. System-centric candidates hand back a list of things that broke: exposure drift, a moved camera, a part arriving rotated . This is the hiring implication behind the whole essay: computer vision assessment has to include the imaging path, because most production vision failures are camera failures that a model-centric interview can never see.
References
- Segment Anything — arXiv (Kirillov et al., Meta AI Research, 2023). (accessed 2026-09-28)
- Announcing the Test of Time Paper Award for NeurIPS 2025 — NeurIPS Blog. (accessed 2026-09-28)
- A3 Vision & Imaging Statistics — Association for Advancing Automation (A3). (accessed 2026-09-28)
- Meta Segment Anything Model 2 — Meta AI Research. (accessed 2026-09-28)
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering — arXiv (Kerbl et al., ACM Transactions on Graphics, 2023). (accessed 2026-09-28)
