@MBGilroyi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
STATEMENT game indeed! @nyliberty showing off that championship DNA out of the gate. Stewie doing Stewie things! One down, eight to go for that 🏆
Michael B. Gilroy retweeted
Impressive, but again - these charts should all be gross profit not ARR
It doesn’t matter if you’re selling dollars for pennies
If AI companies gross margins are 20-30% (high range), these ARR charts are 2-3x overstated
adding Cognition’s growth to a chart that circulated last year (h/t @Yuchenj_UW), you can see that execution speed defines the company
this should not surprise those who saw Scott’s childhood mathcounts video, but hard to fully appreciate until seen in the context of the greats
Michael B. Gilroy retweeted
High ARR multiples won't save an AI startup with runaway burn rates.
Gokul Rajaram @gokulr, Founding Partner at Marathon Management Partners, explains why evaluating companies on revenue alone is a major risk:
"The ARR multiple in isolation without looking at the burn is basically irrelevant…What I want is a company that when the markets turn, they can still keep growing efficiently and they can raise a round even in a bad market."
Using your own brain and instincts have somehow become a competitive advantage. Using AI to *make decisions* in this job is the definition of alpha destruction.
Very useful insight from a portfolio founder: VCs are using Granola and feeding notes into a poorly prompted version of Claude to evaluate and pick apart what founders are saying mid-pitch. The frustrating thing is that AI tends toward groupthink...so "trends" on the internet (whether true or not) color how VCs respond even more today than historically. This founder is tracking those sometimes unfounded trends and turns of phrase and ensuring he rebuts them right out of the gate.
I remember this somehow being the number one signal for business strength during zirp. And then the metric was how quickly and deeply they did a RIF. Short memories.
Michael B. Gilroy retweeted
We’re entering the “$5 Uber” era of consumer - inference edition
VC $ subsidizing huge token usage on products aiming to become indispensable enough to eventually monetize
…and who can stomach the short term burn to try to kill their competitors
Enjoy it while you can!
Michael B. Gilroy retweeted
All the focus on top line growth ignores burn / efficiency.
There are actually two paths, one is super fast growing but highly inefficient / high burn. The other is fast growing but efficient.
If you're in a large market and have grown revenues from 5 to 15M this year, or 3M => 12M, burning sub-5M (burn ratio of 0.5) and you can't raise, please message me.
We at @MarathonMP absolutely love companies that compound efficiently for years.
Kingmaker framework from the product king himself.
The New Kingmakers
It's become clear that revenue scale no longer matters - or is correlated - with raising a round. It's important but not sufficient.
Besides team caliber, what ACTUALLY matters to fundraising is how many top AI-native companies use your products, love your products, embrace your products and profess undying loyalty to you (i.e. say they won't replace it at the drop of a hat).
Why?
AI-native companies are the hardest customers to win. They have the engineers to build it themselves. They have the taste to know when something is mediocre. They try every new tool the week it launches and rip it out the week after. If one of them pays you and stays, your product survived the toughest evaluation there is. Investors know this.
Revenue can be bought. Pilots, services, discounts, a heroic enterprise sales team. ARR built from Fortune 500 pilots proves you can sell. The product might still be mediocre. Loyalty from a kingmaker can't be manufactured.
These companies are the leading indicator. Investors are underwriting where the market will be in 5 years, and the AI-native companies are already living there. Their workflows today are everyone else's workflows in 2029. If they picked you, you've been chosen by the buyers everyone else will copy.
The last reason is distribution. Kingmakers talk. Their engineers post what they use. Their founders compare notes. One of them begets the next three.
Who
There are probably ~50 AI-native companies that are perceived as kingmakers for startups. I've seen $50M+ rounds raised on the backs of just having one of these as a customer. (not pilot, but actual paying customer using the product as a core part of their infra and workflows).
Meanwhile, I see early-stage founders spending 12 months chasing CVS, Walmart, and Pfizer. Long procurement cycles, security reviews, a pilot that never converts. Even when it lands, the logo proves you can survive procurement. That's the wrong thing to prove in 2026.
If you're an early stage enterprise/B2B AI company, Instead of targeting CVS, Walmart, Pfizer, etc, go after the companies on Forbes' AI 50 list.
The new milestone is kingmaker logos. That's what your next round will be priced on.
Replying to @ycombinator
@ycombinator was testing lightweight models for its ai office hours. the goal was useful startup advice at conversational speed.
they moved to glm-5.2 on a dedicated wafer endpoint. the wafer agents tuned the serving setup around their prompts, cache usage, and traffic.
yc then tested it against gpt-4.1 mini on openai and gemma 4 31b on cerebras. wafer had the lowest average llm latency at 379ms.
users on wafer talked to the ai partners for 2.5 minutes longer on average!
read how yc found the right setup
🧵 link in thread
When I consider @immad and the @mercury team I think “product obsessed”. They thoroughly and methodically go to the next highest customer need. Now as an entrepreneur w 10 different excel files sitting in my inbox each week, I’m very excited to onboard this!
1/ Today, we're launching @Mercury Books.
AI-powered accounting software, for your bookkeeper or you, that categorizes and reconciles your transactions the moment they happen.
Michael B. Gilroy retweeted
Look at that latency 😍
As proven by 25+ years of internet services, speed is a feature.
For any app where there is a consumer (or agent!) on the other side waiting for the result, being even a bit faster results in higher conversion (and lower drop-off / churn)
@wafer_ai is the choice for low latency inference 📈🚀
DeepSeek-V4.1-Flash is live on @OpenRouter!!
pick @wafer_ai as your provider
(ss taken 9.12.26)
*overheard at 2nd street blue bottle*
larper 1: "bro Attention is All You Need is one of the most important papers in ai."
larper 2: "yes bro it's so good"
larper 1: "have you read it?"
larper 2: "no, have you?"
larper 1: "no"
don't be like these guys. pls read
we're posting every single resource in the world's most comprehensive ai performance engineering github repo
this is Wafer's ai performance engineering series
resource links in thread 🧵
part 2: "Attention Is All You Need" by Vaswani et al.
arguably the most consequential paper in modern AI.
- the authors cover the original Transformer's computation and information flow.
- scaled dot-product attention, softmax(QKᵀ/√d_k)V, where d_k is the query/key width. the authors scale the logits to counteract large dot products that can produce small softmax gradients.
- separate learned Q/K/V projections for each attention head, followed by per-head weighted aggregation, concatenation, and a learned output projection.
- cross-attention uses decoder queries against encoder keys/values, conditioning predictions on contextualized source representations.
- masking future-position logits before softmax and shifting target embeddings by one position. training parallelizes across known target positions; generation still depends on previous outputs.
- position-wise FFNs transform representations without mixing tokens. the original post-norm architecture applies residual addition before layer normalization.
- sinusoidal positional encodings supply positional information. for any fixed offset, the shifted encoding is a linear function of the original encoding.
- full-sequence self-attention mixing, excluding projections, costs O(n²d), versus O(nd²) for dense recurrence, where n is sequence length and d is representation width. direct connections shorten dependency paths, while all-pairs attention introduces quadratic sequence-length scaling.
- in the authors' English-to-German development experiments, label smoothing improved BLEU despite worse perplexity; varying head count at fixed total attention width did not produce monotonic BLEU gains.
the big model scaled the same encoder-decoder architecture to 16 attention heads, a model width of 1,024, and a feed-forward width of 4,096. the authors report 28.4 BLEU on the WMT 2014 English-to-German newstest2014 test set after 3.5 days on eight NVIDIA P100 GPUs. they averaged the last 20 checkpoints for evaluation. the score exceeded the strongest ensemble listed in Table 2 by over two BLEU.
why this paper was so consequential:
- the authors replaced sequence-aligned recurrence with attention. training could process known sequence positions in parallel, with shorter paths between distant tokens than recurrent models. coauthor Jakob Uszkoreit linked this parallel computation to better use of GPUs and TPUs.
- in 2018, GPT paired a Transformer decoder with generative pretraining; BERT paired a bidirectional Transformer encoder with masked-language-model pretraining. both learned representations from unlabeled text that researchers could reuse across downstream tasks. by October 2020, Google reported using BERT in almost every English query.
- GPT-3's authors evaluated the model using instructions or examples in its context, without task-specific gradient updates. they reported stronger few-shot performance as model size increased. Kaplan et al. quantified empirical relationships between language-model loss, parameters, training data, and compute, and analyzed how to allocate a training budget.
- for the original ChatGPT launch in November 2022, OpenAI adapted a GPT-3.5-family model using dialogue demonstrations and reinforcement learning from human feedback.
- researchers also adapted Transformers to other kinds of data. ViT processed images as sequences of patch embeddings. DiT used Transformers to denoise latent image patches, and the 2024 Sora research model used a diffusion Transformer over spacetime latent patches. these systems used Transformer components for image recognition, image generation, and video generation.
- systems researchers also optimized Transformer training and inference. Megatron-LM partitioned attention and MLP computation across GPUs. FlashAttention used tiling and recomputation to reduce HBM traffic for exact attention. vLLM used PagedAttention so a sequence's KV-cache blocks could occupy noncontiguous GPU memory. Hugging Face Transformers packaged model variants and pretrained weights behind common interfaces. OpenAI's 2020 API launch offered private-beta access to models from the GPT-3 family.
Michael B. Gilroy retweeted
If you see the job of venture capital as generating markups, you get Rivian.
If you see the job of venture capital as generating DPI, you get Tesla.
Total management fees:
Tesla: ~$21M ($20 per $100 invested)
Rivian: ~$115M ($1.44 per $100 invested)
Total carried interest:
Tesla: ~$350M ($333 per $100 invested)
Rivian: ~$300M ($3.75 per $100 invested)
Tesla was the superior investment, by a huge margin, on every metric except two:
- Total venture dollars absorbed
- Total fee income generated
If you see Rivian as a vessel for hot market allocator appetite, the outcome makes sense. It was used and cast aside onto public markets.
This is the fundamental difference between the two companies, and why the industry must avoid valuing capital availability over capital efficiency.
Michael B. Gilroy retweeted
Building a $5B company without 9-9-6
@Immad (Immad Akhund), Co-Founder & CEO, @Mercury, interviewed by @Nakul (Nakul Mandan) @KnuckleupHQ (Knuckle Up with Nakul)
Summary: Akhund has built Mercury to 300,000 customers, a $650M annualized run rate and 4 straight profitable years while refusing almost every intensity norm Silicon Valley reverted to. He is anti-996, still remote-first at 1,200 people, and hires for curiosity and humility. His argument against the grind is arithmetic: 2 extra hours a day never bought him 20% more output, and a sixth employee on a team of 5 does. The line he keeps returning to is that he has engineered his own pressure to fall as the company succeeds, which he says is rare and should be normal.
1. The 9-9-6 Arithmetic. Working 9 to 9 does not make you 20% more productive than working 10 hours. Akhund runs the numbers out loud: the extra 2 hours might buy 5%, and throwing in Saturday might get you to 10%, while going from 5 employees to 6 buys a clean 20%. He also thinks the grind costs you the ideas, because his best ones arrive on a flight or while driving, and a different fork in the road can be worth 10 times more than a few extra hours at the desk. He is not prescriptive about it, and says the tradeoff is reasonable when you are young and cannot afford the extra hire.
2. Culture Is Personalities. Akhund hates the word culture because it is ill-defined, so he reduced it to something measurable: which personalities you hire, and which traits you encourage. Mercury's four founders wrote theirs down at the start, choosing product-minded, humble, helpful and curious, and then built interviews to test against each one. At his previous company he skipped this step, and the salespeople became the loudest voices and set the tone by default. He points at Uber as proof that a high-ego, competitive culture can also build something enormous, so the question is which one fits you and your product.
3. Recruiting Or Nothing. Recruiting only happens when it is the founder's number one priority, and Akhund puts it at about 50% of his time when he is in a hiring phase. He works in phases. He raised $6M from Andreessen Horowitz in 2017, recruited to a team of 8, then stopped hiring for roughly 18 months while they built. Five of those first 8 people had worked with him at his previous company, and all 5 are still at Mercury. After the Series A, advice from Josh at Gusto to hire an in-house recruiter changed the trajectory, and the recruiting team is now around 15 people.
4. The Presentation Interview. Candidates get 45 minutes to prepare a talk on any topic they choose, then present it and defend it. Akhund invented this about 12 years ago as the sales equivalent of an engineering coding challenge, after finding that great salespeople could ace every conventional interview question and then had nothing interesting to say. It grades several things at once: whether they have gone deep on anything, how they handle pushback, and how they respond to a product-thinking curveball like making pickleball popular in America. He says it is one of the best ego detectors he has, because high-ego people cannot take the challenge, and he will never make an exception for a sales hire.
5. No One-On-Ones. Akhund scrapped one-on-ones after hearing Jensen Huang describe running Nvidia without them, and found more CEOs had quietly done the same than he expected. Ten weekly one-on-ones consume a full day plus the head space around it, so he replaced them with group exec meetings: a weekly all-exec, then biweekly GTM, engineering-product-design, compliance and risk, and G&A. Fridays run a 3-hour exec workshop where teams bring work in progress and get decisions made in the room. He says one-on-ones are mostly an avenue for reports to sell themselves. Visibility comes from quarterly written reviews and from watching people perform in group settings.
6. Autonomous Eight-Person Teams. Mercury runs at least 15 product teams of roughly 8 people, typically 1 PM, 1 designer and 6 engineers. Each owns a product surface, a KPI and most of its own roadmap, talks to customers directly, and needs very little approval from Akhund. The friction shows up at the seams. The personal banking team still has to work with risk and disputes, and invoicing still has to work with mobile. With 400 engineers, the whole design is an attempt to recover the efficiency a seed-stage startup gets for free.
7. Metrics Make Local Maxima. Optimizing everything you can measure walks you into a local maximum. Akhund thinks growth teams should be metrics-driven, since sign-up to application conversion has many measurable steps worth grinding on. The counterexample he gives is collecting W-9s from contractors, which materially improves a customer's life and maps to no company-level metric he can name. He points at the Facebook app as a company that became metrics-driven enough to lose its heart. He says EBITDA is the number that finally matters, while most of the route to it resists measurement.
8. Complaining As A Skill. Akhund treats complaining about his own product as a core part of the job. He banks his own fund and his wife's business on Mercury, so he uses it as a customer and files the friction he hits. The skill is manufacturing the mindset of someone who knows nothing about Mercury and gets annoyed at an extra click or a step that makes no sense. He is explicit that product taste at this scale is mostly not his, because the designer on a feature has thought about it far longer than he has.
9. AI Enables The Rep. Mercury tried AI SDRs and Akhund says they were bad then and are probably still bad, because everyone deploying them floods the field and wrecks the signal-to-noise for everybody. What worked was arming human salespeople with context, so an email can open with a specific congratulations on a raise or a note that the company just hired its first finance person. Sales headcount has gone up, and they are hiring aggressively. On support, about 35% of tickets now get answered immediately through Fin, and Akhund expects to bring it in-house to push toward 70% because the next tier requires deep internal integration.
10. The TAM Miscalculation. AI founders keep sizing their market by the labor they replace, and Akhund thinks the number is badly wrong. Once 10 startups compete to automate the same job, the market is whatever the software can charge given that competition, which sits far below the salary line it replaced. The moat question decides where in that range you land. He says this changes how these companies should be valued, and calls it his most controversial view in the conversation.
11. A Regulated Bank That Ships. Mercury has conditional approval to become a chartered bank, and Akhund frames the whole exercise as a contrarian bet. He argues they are already regulated by proxy, running new products and marketing past partner banks, which means third parties dictate the program. The charter buys Zelle, real lending products, direct access to the Fed and later cut-off windows, and removes the disclaimers on mercury.com that cost them trust. He says he would be disappointed if in 5 years they operate exactly like a bank, because the opportunity is building a regulated institution that stays product-led and fast.
12. Bonus Levels. Akhund grew up with parents on welfare and spent his first 10 years as a founder trying to survive, which ended with a $45M exit. The unicorn valuation arrived in 2021, and he describes everything after it as the bonus levels of the game, where the goal turned open-ended and vaguer. Having no scarcity left is what let him commit 3 years to a new bank in 2017, a swing he could not have taken earlier. He has deliberately built things so pressure drops as success climbs, and thinks it is strange that the reverse is normal.
Some will mistake this for reckless investing.
But this is a prepared mind meeting a founder in a moment where high trust & transparency had been built over many interactions. Hundreds of micro datapoints went into the investment decision. Founder friendly meets astute investing.
Neil Mehta agreed to a $500M investment in Rippling in one phone call, on the Friday that SVB collapsed.
"It was entirely possible he was putting money into a sinking business." -- @parkerconrad
Feels like an entire childhood and early adulthood watching this man speak passionate truth on TV. Congrats on an epic career of inspiration and obsession with your craft @RickSantelli
Move fast and break things does not work in financial services. It is not for the faint of heart.