5 Proven Techniques for Token Compression and Prompt Optimization
Reduce costs, improve response quality, and build leaner AI applications with these prompt engineering strategies.

Every token counts. Whether you're building production applications with large language models (LLMs) or running experiments in a notebook, bloated prompts silently drain budgets and degrade response quality. Token compression is the practice of transmitting more intent with fewer tokens, and prompt optimization is how you structure that intent so models respond accurately and efficiently. This guide covers five techniques you can apply right away to reduce token consumption without sacrificing output quality, along with the reasoning behind each approach and practical code examples.
1. Replacing Verbose Instructions with Structured Constraints
Long, conversational system prompts feel natural to write but cost significantly more than tightly structured equivalents. The fix is moving from narrative instructions to declarative constraints, using schema-like formatting that models parse efficiently. Instead of writing:
Please make sure that when you respond, you always use
bullet points and keep answers under 100 words. Do not
include any preamble or sign-off at the end of your reply.
Compress it to:
Format: bullet points | Max: 100 words | Omit: preamble, sign-off
That single line replaces 36 tokens with roughly 14. Across thousands of API calls, the savings compound quickly. Use pipe-delimited key-value pairs, YAML-style constraints, or JSON schema snippets depending on the model family you're working with.
2. Using Few-Shot Examples Strategically, Not Exhaustively
Few-shot prompting — providing example input-output pairs before your actual request — dramatically improves output format consistency. The mistake most practitioners make is adding too many examples. Research from Anthropic and academic benchmarks consistently shows diminishing returns beyond three to five examples for most classification and generation tasks. Here's a lean three-shot prompt for sentiment labeling:
system = """Label sentiment. Reply with one word: Positive, Negative, or Neutral.
Examples:
Input: "Shipped on time and well packaged." -> Positive
Input: "Completely broken out of the box." -> Negative
Input: "It arrived." -> Neutral"""
Three examples establish the pattern. Adding ten more rarely improves accuracy and often introduces contradictions that confuse the model. Audit your existing few-shot prompts and benchmark quality at one, three, and five examples before committing to a larger set.
3. Applying Dynamic Context Trimming for Long Documents
When you pass long documents into a prompt — transcripts, legal text, knowledge base articles — you're almost always paying for tokens the model doesn't need. Dynamic context trimming retrieves only the relevant passage rather than the entire document. Here's a minimal implementation using cosine similarity with sentence embeddings:
from sentence_transformers import SentenceTransformer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
model = SentenceTransformer("all-MiniLM-L6-v2")
def trim_context(query, passages, top_k=3):
q_emb = model.encode([query])
p_embs = model.encode(passages)
scores = cosine_similarity(q_emb, p_embs)[0]
top_idx = np.argsort(scores)[-top_k:][::-1]
return [passages[i] for i in top_idx]
Pass the filtered list of passages instead of the raw document. For a 10,000-token knowledge base where only 800 tokens are relevant, this technique alone can cut context costs by over 90%.
4. Caching Repeated Prompt Prefixes with Prompt Caching
Many applications repeat identical system prompts across every user request: the same persona definition, the same tool descriptions, the same policy constraints. Sending those tokens fresh each time is unnecessary. Several inference providers — including Anthropic with its prompt caching feature and OpenAI with automatic prefix caching — now store and reuse static prompt prefixes server-side, billing cached tokens at a fraction of standard input pricing. Structure your prompts so the stable content comes first and the dynamic content comes last:
[System prompt - static, 800 tokens] <- Cached after first call
[Retrieved context - semi-static, 400 tokens] <- Potentially cached
[User message - dynamic, 50 tokens] <- Always fresh
Before implementing, check your provider's caching documentation. Anthropic's prompt caching kicks in when the cached prefix exceeds a minimum token threshold and the cache is hit within a defined time window.
5. Compressing Chain-of-Thought Reasoning with Scratchpad Separation
Chain-of-thought (CoT) prompting improves model reasoning on complex tasks, but the reasoning trace itself — sometimes hundreds of tokens — often appears verbatim in your API response even when you only need the final answer. That inflates output token costs fast. The fix is to separate the reasoning scratchpad from the final answer using structured output markers:
prompt = """Solve the problem step by step inside tags.
Then provide only your final answer inside tags.
Problem: A warehouse ships 240 units over 6 days at an uneven rate.
Day 1-3 average: 30/day. What is the Day 4-6 average?"""
Your application then parses and discards the <thinking> block, paying for the reasoning tokens but only storing and returning the <answer> content to end users. For APIs that support extended thinking or reasoning modes natively — like Anthropic's extended thinking — the reasoning tokens may be billed at a different rate entirely and can be suppressed from the response body.
Recommended Tools and Resources
- LangChain: Provides token counting utilities and retrieval-augmented generation (RAG) pipelines for dynamic context trimming
- LiteLLM: Unified interface for tracking token usage across providers with built-in cost logging
- Sentence Transformers: Efficient embedding models for semantic passage retrieval
- tiktoken: OpenAI's tokenizer library, useful for pre-flight token counting before API calls
- Anthropic Prompt Engineering Guide: Free documentation covering caching, structured outputs, and CoT best practices
Final Thoughts
Token compression isn't about cutting corners. It's about precision: writing prompts that give models exactly what they need and nothing more. The five techniques above — structured constraints, strategic few-shot sizing, dynamic context trimming, prefix caching, and scratchpad separation — target the most common sources of token waste across production applications. Start by auditing one prompt you use frequently. Measure its current token count, apply one technique, and benchmark the output quality against the original. Incremental, evidence-based optimization is more sustainable than rewriting everything at once. As LLM usage scales, even modest per-call savings translate into real cost reductions and measurably faster response times.
Vinod Chugani is an AI and data science educator who bridges the gap between emerging AI technologies and practical application for working professionals. His focus areas include agentic AI, machine learning applications, and automation workflows. Through his work as a technical mentor and instructor, Vinod has supported data professionals through skill development and career transitions. He brings analytical expertise from quantitative finance to his hands-on teaching approach. His content emphasizes actionable strategies and frameworks that professionals can apply immediately.