Why should `previous_response_id` replace manual conversation stitching in the OpenAI Responses API?
`previous_response_id` can continue a conversation without replaying the whole transcript, but it does not carry every kind of state. This article shows what the API keeps, what it drops, and when you still need your own transcript store.
A second Responses API turn can fail in a subtle way: the reply still looks connected, but the next request no longer has the exact context you thought it did. Replaying the whole transcript avoids that surprise, but it also makes every turn responsible for more state than you may want to resend.
This article explains when `previous_response_id` is the simpler handoff, what the OpenAI docs say it preserves, and which parts of your own state still need to be sent explicitly.
What changes between replaying the transcript and chaining a response
Manual replay and `previous_response_id` solve the same user-visible problem, but they do it through different request shapes. Manual replay sends the prior messages back to the API, so the request body is the conversation record. A chained request sends only the new turn plus the prior response identifier, which lets the platform reconstruct the conversation state from its own stored response record.
That difference matters because the article is not trying to prove that one shape is always shorter. It is showing that the two shapes have different failure modes. Replay is explicit but repetitive. Chaining is compact, but it asks you to understand exactly which state is held by the platform and which state remains your responsibility.
| Strategy | What the request carries | Main benefit | Main limit |
|---|---|---|---|
| Manual replay | The earlier messages and the new turn | The request is self-contained and easy to inspect | The payload grows with every turn and is easy to duplicate incorrectly |
| `previous_response_id` | Only the new turn plus the prior response ID | The next call stays narrow and can reuse the prior response chain | It does not automatically mean every earlier instruction or app-level decision is still present |
SourcesStreaming events and response fields (opens a new tab)OpenAI API reference introduction (opens a new tab)
What the platform carries forward, and what it does not
The important boundary is in the docs: `previous_response_id` links a new request to a previous response, but instructions from the previous response are not carried over automatically. That means a developer cannot assume that a hidden instruction block, a system note, or a prior app decision survives just because the request is chained.
The practical result is that `previous_response_id` is best for continuing a conversation, not for outsourcing all state management. If your app has a policy, role, or guardrail that must remain in effect, send it again. If a piece of context only lives in your database or local memory, keep managing it there instead of assuming the API will infer it from the response link.
const manualReplay = {
instructions: 'You are a support assistant. Keep answers short and remember the refund policy is 30 days.',
input: [
{ role: 'user', content: 'Remember that the refund policy is 30 days.' },
{ role: 'assistant', content: 'Understood. I will keep the 30-day refund policy in mind.' },
{ role: 'user', content: 'Does the policy apply to gifts?' },
],
};
const chainedRequest = {
instructions: 'You are a support assistant. Keep answers short and remember the refund policy is 30 days.',
previous_response_id: 'resp_001',
input: [{ role: 'user', content: 'Does the policy apply to gifts?' }],
}; SourcesStreaming events and response fields (opens a new tab)Data controls in the OpenAI platform (opens a new tab)
Why the local fixture matters even without a live API call
This article does not need a paid model call to prove the decision boundary. The local test only checks the request-shape decision: manual replay carries the full visible transcript, while the chained request carries a prior response ID and a new turn. That is enough to expose the difference the reader needs to reason about before they wire the pattern into production.
The fixture is intentionally small. It models one prior assistant reply and one follow-up user turn because that is the smallest useful case where a chained request looks tempting. The goal is not to claim a benchmark, latency result, or model preference. It is to show the structure that changes and the state that does not disappear just because the request got shorter.
function buildNextRequest(strategy, previousResponse, nextTurn) {
if (strategy === 'manual-replay') {
return {
input: [
...previousResponse.transcript,
nextTurn,
],
};
}
return {
previous_response_id: previousResponse.id,
input: [nextTurn],
};
} When to keep your own transcript store
Keep your own transcript store when the conversation has branches, when you need to edit or redact earlier content, or when your app has state that should survive even if the API chain is reset. Use `previous_response_id` when the main thing you want is a clean follow-up call and the platform-held response chain already contains the context you want to continue.
The decision rule is simple: if the next request should depend on earlier app decisions, keep those decisions in your own data model. If the next request should only continue the prior response chain, `previous_response_id` is usually the narrower and easier choice. That is why the right answer is not 'always use one or the other' but 'treat the response ID as a transport for conversation continuity, not as a replacement for every form of application state.'
- Use manual replay when you need the request body itself to be the source of truth.
- Use `previous_response_id` when you want to continue the conversation chain without resending unchanged text.
- Keep explicit state in your app when instructions, policy, or branching history must remain inspectable and editable.
SourcesStreaming events and response fields (opens a new tab)Data controls in the OpenAI platform (opens a new tab)
Edge cases that break the simple story
The clean rule, 'use `previous_response_id` when you want a shorter follow-up,' breaks as soon as the conversation forks. A branch can share the same prior response identifier, but the two branches can no longer be treated as the same user experience. One branch might be a retry after a timeout, while another branch might be a deliberate change of direction. If your product cares which one happened, the response ID alone is not enough because it only says where the chain came from, not why it diverged.
A second failure case is an instruction update. Suppose your app changes policy after the first turn, or the user edits an earlier requirement. If you chain blindly, the next call can still be anchored to the older response record while your app logic has moved on. In that case, re-sending the relevant instructions is not redundancy; it is how you make the current policy explicit again.
A third boundary is observability. The more you rely on chained responses, the less the next request body tells you by itself. That is good for compactness, but it means diagnostics move into your logs and storage. If you need to replay the exact context later, or hand the transcript to another service, you still need your own persisted record of the messages and the application decisions that surrounded them.
- A branch changes the meaning of the same previous response ID.
- A policy update can make the old chain semantically stale.
- A compact request body can make debugging harder if you do not persist the surrounding app state.
SourcesStreaming events and response fields (opens a new tab)OpenAI API reference introduction (opens a new tab)
A decision rule you can apply before shipping
Use this rule when you are choosing the implementation path. If the next turn should only continue the previous response and you do not need to rewrite, redact, or branch the conversation history, `previous_response_id` is usually the cleaner choice. If the next turn must be reconstructible from the request alone, or if the app owns state that cannot be inferred from the prior response, keep the manual transcript replay.
That rule also keeps the article honest about the boundary of the OpenAI docs. The docs explain the continuation mechanism and the retention window, but they do not remove your responsibility to model business state. The API can reduce how much text you resend; it does not decide what your application should remember, what it should discard, or which parts of the past should still be visible to the next request.
SourcesData controls in the OpenAI platform (opens a new tab)Streaming events and response fields (opens a new tab)
What the local test proves, and what it does not
The Node test is intentionally narrow. It proves that the two request shapes are different and that the instruction text still matters in both shapes. It does not prove model quality, output quality, latency, or how a specific SDK version implements retries. Those claims would need a live service call or a separate integration harness, and this article does not make them.
That limitation is a feature, not a weakness. By keeping the local fixture small, the article avoids turning an API-contract question into a benchmark. Readers who need a production decision still get the useful part: a control-flow difference they can reason about before they write real code. Readers who need operational proof still know exactly which part is unverified and where they would have to measure it themselves.
This is also why the article keeps the transcript replay example beside the chained example. The evidence is not that one is universally better. The evidence is that they answer different questions. One makes the whole conversation explicit in the request. The other makes the conversation continuation explicit in the API call. Once that distinction is clear, the implementation choice is much easier to justify.

