OpenAI recently introduced a powerful feature called Predicted Outputs that can significantly reduce latency in API responses when much of the output content is predictable. Let’s explore this feature through practical examples.
Understanding Predicted Outputs
When modifying text or code files where only small changes are expected, we can provide a prediction of what we think the output will be. The model can then use this prediction to generate responses faster by reusing parts of our prediction that match its intended output.
For more details, see the
Results:
Normal Completion Time: 14,115 ms
Predicted Outputs Time: 4,756 ms
Time Savings: 66%
Total Completion Tokens: 784
Accepted Tokens: 686 (reused, not billed)
Completion Tokens Billed: 98 (only rejected tokens)
Cost Savings: 88%
Tokens per Second (Normal): Approximately 55 tokens/sec
Tokens per Second (Predicted Outputs): Approximately 165 tokens/sec
This demonstrates the power of Predicted Outputs when changes are minimal. Most of the original content was reused, resulting in significant time and cost savings, and a substantial increase in tokens processed per second.
Here is the result in Langfuse from another session recorded shortly after the one above. The difference is still significant.

SOCIAL SHARE CARD GENERATOR