Joined August 2026
Include

Only show posts containing:

Exclude

Hide posts containing:

Time range
-
Minimum likes
Sonnet 5.5 is out, and we're excited about this one. Ahead of its launch, we worked closely with the @AnthropicAI team to test it against the Model ML Composite, our proprietary benchmark for AI in financial services. Several findings stood out to our evaluations team: - Near-Opus quality at a fraction of the cost. Anthropic's latest model stays close to Opus 5.5 in four of our five categories and costs 56–78% less in all five. That puts it right on the cost–performance frontier for analytical finance, PowerPoint creation and multi-document work. - A standout slide builder for the cost. On PowerPoint creation, Sonnet 5.5 comes close to the top slide builders (Opus 5.5, GPT-6 Astra and Fable 5.1) at less than half the cost of any of them. It also beats GPT-5.6 Sol, GPT-6 Sol and Gemini 3.8 Flash on both quality and cost, and it comfortably outscores the best open model, DeepSeek 4.1 Flash, while costing less. - Among the leaders on multi-document work, at a mid-range cost. Sonnet 5.5 sits on the cost–performance frontier, just behind Opus 5.5 and Gemini 3.8 Flash, with GPT-6 Astra leading the category. It also comes in ahead of Fable 5.1 at a fraction of its cost. - On pure-play financial workflows, open models outperform Sonnet 5.5. DeepSeek 4.1 Flash and Kimi K3 both outscore Sonnet 5.5, with DeepSeek at a fraction of the price. Try Claude Sonnet 5.5 in Model ML today.
1
55
It turns out people prefer Sites (HTML dashboards) to PowerPoint and Excel for a lot of use cases…
1
4
62
We just fine-tuned GLM 5.3 and it outperformed GPT-6 Astra and Claude Opus 5 on this task, while costing ~95% less. Here’s what we did. First, we identified a high-volume task where we had a very clear definition of right and wrong. We chose “PowerPoint slide quality checking”: our agent looking at a slide it has built and deciding whether it contains a defect or is clean. We started by running GLM 5.3 out of the box against our test set. It was okay. Precision on defects was 43%, with an F1 score of 52%. Then we built a synthetic training pipeline. We took high-quality slides, programmatically introduced different visual and layout defects, and used frontier models to independently identify and verify the defects before accepting them into the training set. We split the source material before generating corruptions to make sure variations of the same slides couldn’t leak between training and testing. We fine-tuned GLM on 3,244 of these slides, with another 180 held out for validation, then tested it against a completely separate set of 193 slides containing 346 defects. The result was pretty insane. Precision went from 43% → 88%. F1 went from 52% → 77%. That put the fine-tuned GLM ahead of GPT-6 Astra and Claude Opus 5 on the same test set, at ~95% lower cost. Great work by the team - Damiano, Joseph, James and Iñaki.
1
6
99
We're pleased to share that HSBC is working with Model ML to help support adoption of AI across the bank. Teams are using Model ML to assist with research and analysis and to help draft content in HSBC formats and templates. The underlying approach can be configured for different desks, helping colleagues apply AI in a way that fits existing workflows. Model ML is proud of the partnership and grateful to the teams we’ve worked alongside.
1
3
62
Model ML has shipped a lot in the last four weeks. Across everything we built, two themes stood out: enhancing collaboration with your team, and streamlining the work you do together. Here's what went live: - OpenAI feature (10 Aug): We worked with OpenAI behind the scenes on the release of GPT-5.6, running it against Model ML’s Composite, the leading benchmark for AI in financial services. The full paper detailing our work together has been published on their website. - HSBC Asset Management investment (11 Aug): We shared that HSBC Asset Management has invested in Model ML through its flagship VC Strategy. - Image generation (12 Aug): You can now ask the agent to generate or edit an image and it will appear inline in the conversation. Iñaki made this one possible. - Agents sharing skills (14 Aug): Thanks to Simone, agents can now look up skills from your library and share them with each other. - User management (17 Aug): Client admins can provision and deprovision users, promote or demote admins, and export member activity. Great work from Diogo and Dominic. - TresVista (18 Aug): We shared how TresVista scaled their Model ML usage 10x without their costs going up. A huge win from Milan and Vaqar. - Scratch Pad (19 Aug): James brought a new collaborative feature where you and the agent can write, edit and shape together, even mid-run. - Select what the agent edits (24 Aug): Click on specific elements on a PowerPoint slide or Excel sheet and the agent will work on those. Shipped by Mihir. - Document cross-checks (25 Aug): Give the review agent a draft plus the source it should match, and it flags every number and claim that doesn't line up. By Somesh and Diogo. - Setting up teams (26 Aug): Admins can build teams as distribution lists, so work gets shared with the right group. By Diogo. - Word document review agent (27 Aug): Our agent that specialises in document review came to Word, with fixes landing as native tracked changes. Credit to Somesh. - Shared chats (1 Sep): Thanks to Mohit, colleagues can now work with chats you choose to share in a project. - Comments on Sites (2 Sep): Savish and Chester have brought comments to Sites. Get started by clicking anywhere to pin a comment, reply, or mention a teammate. - Financial Technology Partners / FT Partners Firmwide Rollout (3 Sep): I sat down with the great Steve McLaughlin, Founder & CEO of FT Partners to discuss their firmwide rollout of Model ML. - Parallel agents in Excel (4 Sep): Run multiple agent sessions on the same open Excel file. A single workbook no longer means a single task at a time. Shoutout to the Plugins team Isaac and Luca. Much more to come soon as we continue to expand across the globe. Stay tuned.
3
3
226
Next Tuesday, 15 September, Model ML will be running a hands-on workshop at AI Pathfinder's Private Equity AI Strategy Day in London. Tyler Sisk will be joined by Fred Kinsman, Investment Director at Three Hills, to talk about how Model ML and Three Hills have been working together to transform the firm's day-to-day workflows. This will be followed by an interactive session, where they’ll show PE teams how to start doing the same with their own day-to-day work. If you're there, come and find us.
2
55
It's become very obvious that the product is now the harness. But we realized that when we talked to our customers, most people didn’t actually know what an agent harness is, and why it matters. Language models like ChatGPT and Claude take text in, and send text out. That's it. The harness is everything we build around it to turn that into real work. It assembles the context (the information the model needs) sends it to the model, and reads what comes back. When the model wants to act, like creating a PowerPoint deck, all it can do is write text asking to use a tool. The harness reads that request, executes the tool, and actually builds the deck, pulling information from the relevant sources to tell the model what happened. Then it loops, until the work is done. Get the harness right and the same model produces better work on fewer tokens. Watch James, our Head of AI (and by the way, a PhD in Theoretical Physics from @imperialcollege ), explain what an agent harness is, and why it’s necessary, in more detail. Explainer videos about agentic infrastructure are usually a special kind of boring. We tried to make one you'll actually finish. Let us know what you think.
7
108
Claude Fable 5.1 is now live in Model ML. From testing out Fable 5.1 pre-release, a few things stood out to our evals team when they ran it through Model ML’s Composite, the leading benchmark for AI in financial services: - Fable 5.1 displays best-in-class financial reasoning: @AnthropicAI 's latest model leads on financial workflows and is narrowly ahead of Opus 5 and GPT-5.6 Sol on analytical finance. - It’s better at holding context: The clearest gain over Fable 5 is carrying financial context through a multi-step analysis, from connecting the reasoning, to keeping assumptions consistent and evidencing conclusions. Gemini 3.7 Flash remains ahead of Fable 5.1 on multi-document work, and GPT-5.6 Sol leads on complex Excel model construction. On cost, Fable 5.1 is significantly cheaper than Fable 5 for comparable or better performance, but continues to sit at the top end of the cost range for most tasks. Fable 5.1 watermarks its text outputs in line with EU AI Act transparency requirements. This will not have any practical impact on the quality or content of Claude’s outputs. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks. And as with Fable 5, @AnthropicAI 's safety architecture requires 30-day data retention on all Fable 5.1 traffic, so depending on your setup we'll work with your compliance team to determine whether we can switch on access. Try Claude Fable 5.1 in Model ML today.
1
4
238