Prompt Optimization

Prompt optimization is the automated search for the instructions, examples and structure of a language-model prompt, or of a whole multi-step program of prompts, that maximize a measurable score on a set of examples. It treats the text a model is given as a parameter to be tuned by an algorithm rather than written by hand, and it is the main alternative to fine-tuning when the weights cannot or should not change.

Prompt engineering is manual trial and error. Prompt optimization keeps the trial and automates the rest: propose candidate prompts, run them on training examples, score the outputs with a metric, keep what works, repeat. The model's weights never change, so the method works on closed, API-only models. The ingredients are always the same three: a task or program, a metric, and examples. The DSPy documentation notes that the examples can be few, on the order of five or ten to start, and need not be labeled if the metric can judge outputs directly.

Main Approaches

ApproachHow candidates are generatedSource and reported result
Instruction search (APE)A model proposes instruction candidates; the best is selected by scoreZhou et al., November 2022: matched or beat human-written instructions on 19 of 24 tasks
Textual gradients (ProTeGi, TextGrad)A model critiques failures in natural language, and the critique is used to edit the prompt, by analogy with backpropagationPryzant et al., May 2023: up to 31% improvement over the initial prompt. Yuksekgonul et al., June 2024: GPT-4o zero-shot accuracy on Google-Proof QA from 51% to 55%
LLM as optimizer (OPRO)The optimizer model sees past prompts with their scores and writes better onesYang et al., September 2023: up to 8% over human-designed prompts on GSM8K, up to 50% on Big-Bench Hard tasks
Program compilation (DSPy)Pipelines are declared as modules; a compiler bootstraps demonstrations and instructions for eachKhattab et al., October 2023
Reflective evolution (GEPA)A model reflects on full execution traces and candidates are kept along a Pareto frontierAgrawal et al., July 2025

These figures come from each paper's own tasks and models and are not comparable across rows.

DSPy and Its Optimizers

DSPy, from Stanford, is the most widely used framework. Its premise is that prompts should not be strings in application code: the developer declares what each step takes and returns, composes steps into a program, and an optimizer fills in the wording. This matters most for multi-stage pipelines, where a metric exists only on the final output and credit must be assigned across steps.

As of October 2026 the DSPy documentation lists several optimizer families. Few-shot optimizers such as BootstrapFewShot run the program, keep the traces that pass the metric and reuse them as demonstrations. MIPROv2 proposes instructions grounded in the program's code, data and traces, then uses Bayesian optimization over combinations of instructions and demonstrations; the underlying MIPRO paper (Opsahl-Ong et al., June 2024) reported beating baseline optimizers on five of seven multi-stage programs, by up to 13% accuracy, using Llama-3-8B. GEPA reflects on trajectories, including reasoning and tool calls, to diagnose failures and propose prompt updates, and can use domain-specific textual feedback. Its paper reported outperforming MIPROv2 by more than 10% and, across six tasks, outperforming GRPO reinforcement learning by 6% on average while using up to 35 times fewer rollouts. That comparison comes from GEPA's authors and has not, to this page's knowledge, been independently replicated at scale.

The documentation puts a typical simple optimization run at about two US dollars and ten minutes, with a range from cents to tens of dollars depending on model and data.

When It Beats Fine-Tuning, and When It Does Not

Favors prompt optimization: the model is available only through an API; there are tens of examples rather than thousands; the base model will be replaced in months, and a prompt can be re-optimized against the new one far faster than a dataset can be retrained; the system is a pipeline of several calls; or a human needs to read and audit what changed. An optimized prompt is inspectable text. A weight update is not.

Favors fine-tuning: the task needs knowledge or skill the model lacks and cannot be handed in context; latency and cost per call matter, since long optimized prompts with many demonstrations are paid for on every request; or the goal is to move a task onto a small model. Liu et al. (May 2022) argued that few-shot parameter-efficient fine-tuning was both more accurate and cheaper than in-context examples, though on models from that era.

Often both. The BetterTogether study (Soylu et al., July 2024) alternated prompt optimization and weight fine-tuning on the same pipeline and reported gains over optimizing weights alone and prompts alone of up to 60% and 6% respectively, on small open models and three task types. DSPy ships this as an optimizer, along with BootstrapFinetune, which distills an optimized prompt program into weight updates.

Limits

An optimizer maximizes the metric it is given. With small development sets it overfits, and the DSPy documentation suggests 200 or more examples for longer MIPROv2 runs for that reason. With a flawed metric it finds the flaw, the prompt-level counterpart of reward hacking; metrics built on LLM judges are particularly exposed. Optimized prompts are also specific to the model they were tuned on and may not transfer. Prompt optimization is best understood as one layer of context engineering: it tunes instructions and examples, while retrieval, tools and memory decide what else the model sees.