Natural language processing (NLP) is the branch of machine learning focused on text and speech: reading language, representing it numerically, and producing labels, spans, translations, summaries, or answers. Inputs are usually sentences or documents. A tokenizer breaks them into tokens, models build representations from those tokens, and a task head or generative decoder produces the output you measure.
Take a simple sentence: “This definition is so informative.” A common tokenizer splits it into tokens such as This, definition, is, so, and informative. Those units are not always whole words. Subword tokenizers may split rare words further so the model can handle vocabulary it never saw at training time. Everything downstream, from BERT-style classifiers to large language models, starts here.
If you ship anything that reads user text, you are doing NLP whether or not the product label says so. Search queries, support tickets, chat messages, and document uploads all pass through the same primitives. The engineering work is choosing representations, picking or fine-tuning models, and monitoring whether behavior holds in production.
Build better agents with Arize
Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.
Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability.
Key takeaways
- NLP turns raw language into tokens and numerical representations, then maps those representations to a task output such as a class label, span, or generated string.
- Tokenization choices affect cost, context limits, and multilingual behavior. They are not a cosmetic preprocessing step.
- Classification was the default NLP product shape for years. Modern LLM applications often generate text, but classification, routing, and scoring still show up inside pipelines.
- Pre-trained transformers changed the workflow: fine-tune a general model on a narrow task instead of training feature engineering stacks from scratch.
- Production NLP needs evaluation and monitoring on real inputs, not only offline accuracy on a static test set.
From tokens to representations
Tokenization is the first contract your system makes with the data. Subword and byte-level tokenizers trade human readability for coverage on typos, rare terms, code, and mixed language. After tokenization, models embed tokens and mix context across the sequence. Encoders such as BERT emit pooled or per-token vectors. Decoder-only LLMs predict the next token. Representation quality sets the ceiling for every task on top, which is why teams still study transformer behavior even when the product surface is a chat box.
For a walkthrough of how BERT popularized the transformer encoder pattern for NLP tasks, see Unleashing BERT for NLP.
Classification and structured prediction
For much of the 2010s, the most common NLP deployment was classification: sentiment, intent, topic, toxicity, language ID, spam. You attach a linear head to a pooled embedding and train on labeled examples. Inference returns a discrete label and often a confidence score.
Named entity recognition, sequence labeling, and relation extraction still appear inside LLM systems as routers, safety filters, and extractors. Monitoring classification in production means tracking class distributions, calibration, and error slices by locale or channel. See NLP sentiment classification monitoring for a concrete starting point.
Where LLMs fit in the NLP picture
Large language models blur the line between “understand” and “generate.” Many tasks that used to require a dedicated classifier can be prompted or fine-tuned in a single model: extract JSON fields, classify intent, rewrite text, answer questions over a passage. The NLP stack becomes orchestration around one general model plus tools, retrieval, and guardrails.
Classic NLP concerns remain. Context windows are measured in tokens. Long documents need chunking. RAG depends on embeddings. Tool calls need parsers that fail when output format drifts. A fluent generative answer can still be wrong, and a classifier accurate on average can fail on one segment.
Ship with task-specific metrics, human review on slices, and traces linking inputs to outputs. What AI engineering covers includes data, evaluation, and observability alongside model choice. Log tokenizer settings, model version, and latencies. Watch for train-serve skew, label schema drift, tokenization surprises on SKUs or codes, and metrics that hide rare high-value failures. NLP without monitoring is a lab result, not a system.
FAQ
Is NLP the same as large language models?
No. LLMs are one family of NLP models, heavily used today. NLP also covers tokenizers, classical features, smaller encoders, speech pipelines, and task-specific heads. Many products combine an LLM with smaller NLP components.
Why do token counts differ between models?
Each model ships with its own tokenizer and vocabulary. The same English sentence can map to different token counts across providers, which changes cost and how much text fits in context.
Do I still need a separate classifier if I have an LLM?
Sometimes no, sometimes yes. LLMs can classify via prompting. Dedicated classifiers can be cheaper, faster, or easier to monitor at high volume. Routing, safety, and compliance layers often stay as explicit classifiers even when generation uses an LLM.
What is the first metric to monitor for a text classifier?
Start with prediction distribution over classes plus accuracy or F1 on a frozen golden set replayed daily. Add slice metrics once you know which cohorts matter commercially.
How does NLP relate to RAG and agents?
RAG uses NLP embeddings and often NLP chunking to retrieve text. Agents use NLP to parse user messages, tool outputs, and documents. The generative step is NLP too. Evaluation spans retrieval quality, tool use, and final language output.