The GLM-5.3 Flash model ranks first among open-weight AI models on the Finance Agent v2 benchmark, placing fifth overall and ahead of Claude Fable 5. At $0.048 per test, the model is the cheapest option near the top of that ranking and costs roughly 18 times less than its predecessor, GLM 5.3. Baseten released the system, also referred to as GLM-5.3 Fast, to serve real-time use cases requiring consistent performance.

This is the first model in the GLM suite to accept images, scoring 86.0% on the MMMU benchmark to rank fourth among open-weight models. It features a 1 million token context window, supports tool calling, and ranks eighth among open-weights on the Vals Index. Tests for the model, codenamed ox-alpha, were run at temperature 1 and top_p 0.95 via Fireworks AI and a native API for output limits of up to 128,000 tokens.

Sign in to suggest edits

Key sources

  1. SOURCE@baseten“Designed for real-time use cases that demand consistent performance”x.com
  2. SUPPORT@valsai“costs roughly 18x less than GLM 5.3”x.com
  3. SUPPORT@valsai“$0.048 per test it is also the cheapest model anywhere near the top of that board”x.com
  4. SUPPORT@valsai“the model has a 1M-token context window and supports tool calling”x.com
Markdown