---
format: "aidr-story-markdown/v1"
id: "a0ecdfee72fef23879c0259187a975b33cf25b133b38cee93a51b8d28f8addca"
canonical_url: "https://aidr.today/a0ecdfee?lang=en"
title: "Optimizing cost and latency with Amazon Bedrock prompt caching"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-15T16:18:19.000Z"
category: "Infra"
topics: ["amazon","agent","multi-agent","inference","benchmark"]
source_urls: ["https://aws.amazon.com/blogs/machine-learning/optimizing-cost-and-latency-with-amazon-bedrock-prompt-caching/","https://aws.amazon.com/blogs/machine-learning/optimizing-agent-system-prompts-with-amazon-bedrock-agentcore/"]
summary: "Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration."
---

# Optimizing cost and latency with Amazon Bedrock prompt caching

> [Open the canonical story](<https://aidr.today/a0ecdfee?lang=en>)

**Published:** 2026-09-15T16:18:19.000Z
**Category:** Infra
**Topics:** amazon, agent, multi\-agent, inference, benchmark

## Summary

Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models\. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration\.

## Sources

- [Story source](<https://aws.amazon.com/blogs/machine-learning/optimizing-cost-and-latency-with-amazon-bedrock-prompt-caching/>)
- [Story source](<https://aws.amazon.com/blogs/machine-learning/optimizing-agent-system-prompts-with-amazon-bedrock-agentcore/>)

