---
format: "aidr-story-markdown/v1"
id: "8554e7ebf3c683eb16e0e320179e779f4691a062558c5eea0bac93abe8dca186"
canonical_url: "https://aidr.today/8554e7eb?lang=en"
title: "OpenAI Cuts Long Chat Load Times 94% in Largest Performance Upgrade for Codex"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-15T08:44:18.000Z"
category: "Releases"
topics: ["OpenAI","performance","Codex","optimization"]
source_urls: ["https://huggingnews.com/ai/openai-cuts-long-chat-load-times-94percent-in-largest-performance-upgrad-0b585c9e","https://x.com/scaling01/status/2088474114382025096","https://x.com/i_dg23/status/2088515880032301083","https://x.com/kimmonismus/status/2088529353722270201"]
summary: "OpenAI will launch an infrastructure overhaul for ChatGPT and Codex next week that allows the software to load only the necessary segments of conversations instead of rendering entire history logs. In an internal test of a 741 turn, 231 MB thread, the average load time dropped from 27.6 seconds to 1.66 seconds. The upgrade reduces overall memory growth by 41% and limits the number of required requests to 16, down from 894 in the previous version. The update cuts the number of transcript items loaded from 15,529 to 64 per session to reduce stuttering and memory pressure during scrolling. This architecture creates the technical groundwork for AI agents that can operate continuously over long periods."
---

# OpenAI Cuts Long Chat Load Times 94% in Largest Performance Upgrade for Codex

> [Open the canonical story](<https://aidr.today/8554e7eb?lang=en>)

**Published:** 2026-08-15T08:44:18.000Z
**Category:** Releases
**Topics:** OpenAI, performance, Codex, optimization

## Summary

OpenAI will launch an infrastructure overhaul for ChatGPT and Codex next week that allows the software to load only the necessary segments of conversations instead of rendering entire history logs\. In an internal test of a 741 turn, 231 MB thread, the average load time dropped from 27\.6 seconds to 1\.66 seconds\. The upgrade reduces overall memory growth by 41% and limits the number of required requests to 16, down from 894 in the previous version\. The update cuts the number of transcript items loaded from 15,529 to 64 per session to reduce stuttering and memory pressure during scrolling\. This architecture creates the technical groundwork for AI agents that can operate continuously over long periods\.

## Sources

- [Story source](<https://huggingnews.com/ai/openai-cuts-long-chat-load-times-94percent-in-largest-performance-upgrad-0b585c9e>)
- [Story source](<https://x.com/scaling01/status/2088474114382025096>)
- [Supporting source](<https://x.com/i_dg23/status/2088515880032301083>)
- [Supporting source](<https://x.com/kimmonismus/status/2088529353722270201>)

