Natural language processing (NLP) is the discipline of building systems that read, transcribe, translate, and answer human language: the classifiers, extractors, parsers, and language models that sit upstream of search, assistants, and translation services. Its older crafts have not gone away; they have industrialized. The WMT 2025 general machine translation task drew a record 36 teams, and its organizers still route the final ranking through human evaluation because automatic metrics cannot be trusted on their own . That pattern, old techniques under new scale, defines hiring in this field: the scarce candidates are the ones who hold both the classic machinery and the modern model stack, and most CVs only show one of the two.
Challenges in Natural Language Processing Recruiting
Machine translation still runs a human evaluation economy
Machine translation looks solved from the outside and refuses to be solved from the inside. WMT 2025 drew a record 36 unique teams, and the preliminary ranking paper is explicit that the official standings will come from human evaluation, which supersedes the automatic scores . The reason sits in the metrics task: current automatic evaluation systems still show major weaknesses, including susceptibility to fluent but semantically irrelevant output, systematic gender bias, and poor correlation with human judgments on low-resource languages . The hiring lesson follows directly. Translation experience means owning evaluation, human or learned, in the loop. A candidate who has shipped machine translation can discuss reference-free scoring, challenge sets, and the specific pair where their metric lied. A candidate who has only called a translation API cannot, and the interview that does not test the difference hires an API user for an evaluation job.
Information extraction remains the pipe that feeds every product
Information extraction did not vanish when language models arrived; it moved underneath them. Every retrieval stack, knowledge base, and document product still needs entities pulled out of unstructured text, relations typed, values normalized, and schemas enforced. The LLM makes that easier to prototype and harder to verify, because a model that extracts quietly rewrites the sentence while claiming fidelity. The split among candidates runs between people who treat extraction as a prompt and people who treat it as a pipeline with provenance: canonical forms, constraints, and adjudication across fields. The second population is smaller and far harder to find, and it is the population a document-heavy product actually needs. The probe is cheap: hand over two overlapping extraction schemas and ask what happens to a value that satisfies neither.
Sentiment analysis fragments into domain dialects
Sentiment analysis is the clearest case of a task that never generalized. General models read financial text badly because the vocabulary is specialized, which is why FinBERT's 2019 work continued pretraining BERT on financial corpora and improved the state of the art on classification accuracy by 15% on its benchmark datasets . The pattern repeats in every jargon-heavy domain: pharma, energy, legal, clinical. A candidate with product-review sentiment experience and a candidate with earnings-call sentiment experience share a library and almost nothing else, because the neutral class means different things and the label schemas disagree. Hiring for sentiment analysis means naming the domain and testing against its dialect, not hiring for the task in the abstract, and a generalist score from a general model proves neither one can read the domain's meaning.
Speech processing now shares a benchmark with reasoning models
Speech processing moved from a niche into the standard benchmark suite. MLPerf Inference v5.1 added Whisper Large V3, paired with a modified Librispeech dataset, as one of three new benchmarks alongside DeepSeek-R1 and Llama 3.1 8B, with support for both datacenter and edge systems . That placement says what the market already knew: transcription, speaker work, and speech interfaces are mainstream infrastructure with mainstream latency budgets. The candidate pool has not caught up. ASR engineers think in acoustic features, streaming constraints, and word error rates; LLM engineers think in tokens and prompts. Speech products need both in one head, or two heads that talk to each other, and the title rarely says which one is responding in an interview.
Conversational AI moved onto the voice channel
Conversational AI now carries a phone call, and the economics are industrial. Salesforce pegs the call center market at more than $135 billion annually, with each customer-service resolution costing about $5 or more, and describes its Agentforce Voice work as moving beyond what chatbots could do: listening, interpreting, and escalating calls with context . Voice changes every constraint at once: latency budgets in hundreds of milliseconds, interruptions, turn-taking, silence handling, and ASR errors feeding dialogue state. A candidate who has built text chatbots and one who has built voice agents differ on barge-in and end-of-speech detection alone. Teams hiring conversational AI for the voice channel need evidence of live-call work; teams that accept text-demo evidence are hiring for a channel they are not shipping.
Text analysis skills hide under every other NLP title
Text analysis is the craft underneath the headlines: tokenization, morphology, parsing, corpus work, annotation design, and the quiet decisions about what counts as a word. Those skills rarely appear as a job title, but every serious natural language processing team eventually rebuilds a tokenizer, fixes a normalization bug, or re-annotates a dataset when a benchmark stops measuring the right thing. Candidates who came up through modern APIs often skipped the layer entirely, which is fine until a compounding language breaks the tokenizer and nobody can explain why throughput doubled. The hiring signal is old-fashioned and reliable: people who have done text analysis talk about annotation agreement and tokenization edge cases without being asked, and most modern-stack candidates cannot, which is why the skill hides inside every other title on the team.
Natural language processing (NLP) claims fail under a language-specific probe
Verification in this discipline is language-specific by definition. A candidate claiming natural language processing (NLP) depth should be able to walk through a task in a real language pair or domain: the dataset, the annotation scheme, the metric, and the case where it failed. The WMT25 metrics findings supply the trapdoors to use: fluent but semantically irrelevant output, gender bias, collapsed correlation on low-resource languages . Asking how a candidate's system behaved on a low-resource pair, or how they validated a translation in a language they do not speak, separates operators from API callers within two questions . This is the hiring implication behind every challenge above: NLP assessment has to run inside the target language reality, because a candidate who can only narrate English demo behavior cannot own a multilingual or domain-shifted product.
References
- Preliminary Ranking of WMT25 General Machine Translation Systems — arXiv (WMT 2025 General Translation Shared Task). (accessed 2026-09-28)
- Findings of the WMT25 Shared Task on Automated Translation Evaluation Systems — Association for Computational Linguistics. (accessed 2026-09-28)
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models — arXiv (Araci, University of Amsterdam, 2019). (accessed 2026-09-28)
- MLCommons Releases New MLPerf Inference v5.1 Benchmark Results — MLCommons. (accessed 2026-09-28)
- How Voice AI Is Reshaping the $135 Billion Call Center Industry — Salesforce News. (accessed 2026-09-28)
