tinkering with LLMs

Portland
Joined April 2013
my days are like `N` hours of working with frontier models, babysitting them in frustration and disbelief at the frequent errors and slop interrupted by `M` min breaks of browsing X takes about how the same models (and _especially_ the next gen!) are superhuman at ~everything
48
118
19
1,531
80,233
David Huang retweeted
I got this question so many times today. "How can you grow 80% at $7B?" The true answer is that we're finally seeing a breakthrough with AI agents starting to work in the enterprise. The AIs have been super smart for a while, but have lacked basic context that's in people's heads, or in some SaaS system-or-record. A lot of organizations are deploying FDEs to capture this context, or Ontology, and feed it to the AI. This is labor intensive and expensive. We just automated that with Genie Ontology. Once you have that enterprise context graph, an AI agent like Genie becomes magical. I find myself no longer waiting for answers from my CRO, CFO, CMO, CHRO etc, I just keep queuing up questions on the phone while sitting in meetings. It'd frankly addictive. Our customers are starting to do the same, over 70% of all queries on the platform are now generated by Genie agents. This fuels more questions to the platform, which drives consumption, which drives revenue. That's the simple answer.
79
220
35
1,233
246,466
David Huang retweeted
We worked with @SpaceXAI to evaluate Grok 4.6 on the latest OfficeQA Pro V2 from @DbrxMosaicAI. It achieves the SOTA performance with @databricks's Genie harness! The model is strongest on our document understanding and data reasoning tasks, and it is a very efficient driver!
66
87
42
746
3,470,639
David Huang retweeted
Really excited to open source a new project: Omnigent, a meta-harness for AI agents. It lets you build multi-agent coding and custom agents, sitting above Claude Code, Codex, Pi, and agent SDKs to let you compose them. It also adds live collaboration and rich control policies.
97
214
62
1,279
244,575