@hexiangi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
@SpaceXAI: @grok multimodal, imagine | prev @GoogleDeepMind: gemini & imagen
Earth
Joined September 2014
- Tweets1K
- Following557
- Followers5.9K
- Likes3.6K
Pinned Tweet
Thrilled to receive this feedback from Tesla — they tackle the world’s most critical multimodal challenges.
Extremely proud that the team’s work is paying off🚀
Grok 4.6 multimodal is a step change from Grok 4.5. It’s one of the under-discussed improvements, and I’ve been very impressed by it.
My daily work includes reviewing lots of videos and understanding the context; Grok 4.6 improves the productivity of such workloads by at least 10x if not 100x.
Such workflows may not be captured by common VLM benchmarks, but in my use cases it outperforms Gemini 3.5 and Gemma 4, which is considered the SOTA of VLMs in my opinion.
Hats off to the multimodal teams—you did a great job.
Really looking forward to seeing first gemini response from the space 🚀
Very cool to work with @PJaccetturo and his team at Genre AI on this pilot scene of the Odyssey.
If you have any questions about how to use Grok @imagine for cinematic workflows, drop them below.
We're always looking to improve, let us know what tools you're missing or ideas you have 🙏
Hexiang (Frank) Hu retweeted
xAI's Grok Imagine Image 2.0 takes #4 on the Artificial Analysis Text to Image Leaderboard, the highest ranked model outside OpenAI, a 14 place climb over the previous generation, and lands on the Pareto frontier for quality vs price
Grok Imagine Image 2.0 is xAI's latest image model released in August. The model supports text to image, image editing, as well as multi-reference generations with up to five input images.
Grok Imagine Image 2.0 is a large step up on the previous generation grok-imagine-image-quality: it climbs from #18 to #4 in Text to Image and from #16 to #10 in Image Editing, closing the gap to the leaders on every Text to Image capability and use case we measure, with the largest gains in Knowledge, Text Rendering and Lighting.
Grok Imagine Image 2.0 is available through the xAI API as grok-imagine-image-2.0, in Grok Imagine on grok.com and the Grok apps, and through fal and Replicate.
Congratulations to @SpaceXAI and @elonmusk on the release!
See below for our analysis and example outputs of Grok Imagine Image 2.0 in the Artificial Analysis Image Arena 🧵
Hexiang (Frank) Hu retweeted
building models that understand us requires building models of us - super excited to share some work in that direction and hope people find it useful!
Huge congrats on the new model launch @ericzelikman @YuchenHe07 🚀💫🌟
LFG
Image + video + smart agents work better together at grok.com/imagine/agent
A solid step in agentic video creation — we’d love your feedback!
Looped transformers aren’t a new species. They’re a way to buy FLOPs without buying parameters.
If intelligence tracks useful compute more than weight count, that’s the correct trade.
Hexiang (Frank) Hu retweeted
Replying to @heyruchir
This is why companies that don’t pull the plug on you matter: OpenAI will cut Cursor after the acquisition, then keep using X to promote their models.
The asymmetry is the whole story.
Hexiang (Frank) Hu retweeted
Replying to @juliarturc
Providers start an API business before they know which apps will matter. Cursor was one of those apps, used the API well, and the provider made a lot of money off it.
Asking why they “allowed” Cursor from day one gets the sequence backward. First comes the API, then the apps, then the scale.
The scary part isn’t Cursor. It’s the precedent: downstream startups now have to compete with the market and with their own provider (OAI in this case) flipping on them.
Hexiang (Frank) Hu retweeted
Grok Bot is now available to everyone with a standard Grok or Cursor subscription. It's grown faster than any product we've seen.
It's been particularly exciting to see the range of jobs people delegate to Grok Bot, from running small e-commerce businesses (including support, advertising, inventory, finance), coordinating customer events (directly pinging and working with dozens of human coworkers), testing production software, and completing large, mundane parts of users' day-to-day work.
Hexiang (Frank) Hu retweeted
I gave Grok @bot a link to the post below and asked it to take up the cinematography challenge.
Direct the film entirely through the chat casually and nothing else. Pretty wild.
Homer had a lyre. You have Grok Imagine.
Create a compelling scene from Homer’s The Odyssey that shows what Grok @Imagine’s video and voice capabilities can do. We’re awarding $100K, $50K, and $25K to the top three videos submitted by quoting this post. 🧵
Imagine 2 is designed for real-world usefulness.
This is it working in the hands of art professionals — assets rebuilt from screenshots, already usable.
Thanks Dogan for sharing this feedback over♥️
Just got this text from an old friend.
His exact words:
"Grok Imagine turned out really good, man. It saved us from ChatGPT. I gave it two screenshots from our game, and it went through them one by one, recreated all the items, cropped them, and gave me EVERYTHING export-ready!"
He’s a senior art director on a team that actively develops games.
It was quite a nice moment.
Hexiang (Frank) Hu retweeted
I asked Grok @bot to explain rank, spectral power, and optimization in geometrical terms, and compare them with Adams, Muons, and Aurora. Below is the result.
If you have any suggestions on how to make this video better and more intuitive, please comment below. My Grok @bot will pick them up and incorporate your ideas.