{"slug":"baml","name":"BAML","domain":"boundaryml.com","verdict":"As of 2026-07-13, ChatGPT, Claude, Gemini, Grok collectively rank BAML #3 of 10 for structured output tool for llms (one of 2 leaderboards it appears on). Source: https://modelsagree.com/product/baml (modelsagree.com, CC BY 4.0).","best_rank":3,"categories":2,"brief":{"category":"best-llm-structured-output-tool","title":"Best structured output tool for LLMs","rank":3,"of":10,"top":"Instructor","day":"2026-07-17","why":[{"t":"contract-first type-safe client code","m":["ChatGPT","Gemini","Claude"],"q":"contract-first design, and robust parsing"},{"t":"resilient Schema-Aligned Parsing","m":["ChatGPT","Gemini","Claude"],"q":"Its Schema-Aligned Parsing is incredibly resilient"},{"t":"cross-language, cross-provider option","m":["ChatGPT","Gemini","Claude"],"q":"the best cross-language, cross-provider option"},{"t":"first-class streaming of partial typed objects","m":["ChatGPT","Claude"],"q":"first-class streaming of partial typed objects"}],"gap":[{"t":"thin library instead of a framework","m":["Claude"],"q":"stays a thin library instead of a framework"},{"t":"seamless Pydantic integration","m":["Claude","Grok","ChatGPT","Gemini"],"q":"Dominant Python library with seamless Pydantic integration"},{"t":"almost zero learning curve","m":["Gemini"],"q":"almost zero learning curve"}],"fix":[{"t":"proprietary DSL and build step","m":["ChatGPT","Claude","Gemini"],"q":"You must adopt a proprietary DSL and build step"},{"t":"overkill for simple projects","m":["Claude","Gemini"],"q":"over-engineered for simple or Python-only projects"},{"t":"teams wanting plain-library ergonomics","m":["ChatGPT","Claude"],"q":"a harder sell to teams wanting plain-library ergonomics"}]},"entries":[{"slug":"best-llm-structured-output-tool","title":"Best structured output tool for LLMs","rank":3,"of":10,"score":12,"appearances":3,"modelRanks":{"ChatGPT":1,"Claude":4,"Gemini":1},"reason":"Schema-first contracts, generated type-safe clients, resilient schema-aligned parsing, streaming, and built-in retry/fallback logic across model providers make it the strongest production package","reasons":[{"model":"ChatGPT","reason":"Schema-first contracts, generated type-safe clients, resilient schema-aligned parsing, streaming, and built-in retry/fallback logic across model providers make it the strongest production package"},{"model":"Gemini","reason":"(BAML is in a near-tie with Instructor but takes the top spot due to its multi-language compilation, contract-first design, and robust parsing.) It decouples prompts and schemas into dedicated .baml files that compile to type-safe client code in Python, TypeScript, Go, and Rust. Its Schema-Aligned Parsing is incredibly resilient, recovering data from malformed LLM outputs to prevent costly retry API calls."},{"model":"Claude","reason":"Its Schema-Aligned Parsing repairs malformed output (markdown fences, trailing commas, truncated JSON) instead of constraining decoding, so it extracts reliable structured data from any model — including small local ones with no native JSON mode; typed prompt functions with codegen for Python/TypeScript/Ruby and beyond, first-class streaming of partial typed objects, and an integrated test playground make it the best cross-language, cross-provider option."}],"fixes":[{"model":"ChatGPT","fix":"Remove the separate DSL/code-generation workflow by supporting equally robust native-language schemas"},{"model":"Claude","fix":"You must adopt a proprietary DSL and build step — a real toolchain commitment that is overkill for a single extraction endpoint, and a harder sell to teams wanting plain-library ergonomics."},{"model":"Gemini","fix":"It forces teams to adopt a proprietary domain-specific language (DSL) and integrate its custom compiler toolchain, making it over-engineered for simple or Python-only projects."}],"updated":"2026-07-13","rank_history":{"days":["2026-07-12","2026-07-13"],"ranks":[2,2]},"reasoning_shift":[{"model":"Gemini","from":"2026-07-12","to":"2026-07-13","added":[{"t":"contract-first design","q":"contract-first design"},{"t":"Rust type-safe client code","q":"type-safe client code in Python, TypeScript, Go, and Rust"},{"t":"proprietary DSL and custom compiler","q":"It forces teams to adopt a proprietary domain-specific language (DSL) and integrate its custom compiler toolchain"}],"dropped":[{"t":"direct imports of native classes","q":"Support direct imports of native Python classes or TypeScript interfaces"}]},{"model":"Claude","from":"2026-07-12","to":"2026-07-13","added":[{"t":"small local models","q":"including small local ones with no native JSON mode"},{"t":"streaming partial typed objects","q":"first-class streaming of partial typed objects"}],"dropped":[]}],"api":"https://modelsagree.com/api/v1/best/best-llm-structured-output-tool.json"},{"slug":"best-prompt-engineering-framework","title":"Best prompt engineering framework","rank":6,"of":14,"score":3,"appearances":1,"modelRanks":{"Claude":3},"reason":"Prompts as typed functions in a purpose-built DSL with first-class tests, a VS Code playground, and Schema-Aligned Parsing that recovers usable structure from imperfect model output (more forgiving than strict JSON mode); brings version control, types, and unit testing to prompt engineering.","reasons":[{"model":"Claude","reason":"Prompts as typed functions in a purpose-built DSL with first-class tests, a VS Code playground, and Schema-Aligned Parsing that recovers usable structure from imperfect model output (more forgiving than strict JSON mode); brings version control, types, and unit testing to prompt engineering."}],"fixes":[{"model":"Claude","fix":"Requires adopting a new DSL and codegen build step and buying the whole team in; smaller ecosystem and fewer integrations than the incumbents."}],"updated":"2026-07-14","api":"https://modelsagree.com/api/v1/best/best-prompt-engineering-framework.json"}],"page":"https://modelsagree.com/product/baml","check":"https://modelsagree.com/check?q=BAML","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}