@nm_micro1i
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Joined August 2024
- Tweets84
- Following15
- Followers15
- Likes129
Neet M retweeted
we are opening up our first robotics data lab in Malibu!
in 9 months, our robotics department has grown from 0 to $100m ARR.
this data lab is an effort to accelerate all the progress being made, with an increased focus on evaluating models with hardware.
we are hiring robotics researchers, engineers, and teleoperators.
please reach out if you are interested in training robots at the beach. 🏄♂️
Neet M retweeted
Today we’re launching micro1’s PII transformation model, flow-transform 1.0, delivering frontier-level performance across detection, identity synthesis, and transformation of personally identifiable information.
On PrivacyBench, our model reaches 96.0% F1, outperforming every detection baseline we tested, including Tonic Textual, Claude Opus 4.8, Sonnet 4.6, Microsoft Presidio, Haiku 4.5 and GLiNER2.
Some of the most valuable training data for frontier AI models lives inside fully functioning companies. It captures years of real work across decisions, communications, tools, handoffs, exceptions and the relationships connecting them.
The problem is that this data is also full of PII.
Traditional redaction makes the data safe, but it also destroys the very workflows and relationships frontier models need to learn from.
flow-transform 1.0 solves this by turning enterprise operational data into high-fidelity training data for frontier models by replacing real-world identities without flattening the reality the data captures.
Neet M retweeted
We’re sharing a deeper look at flow, our next-generation data platform for delivering predictable units of intelligence improvement at scale.
This article transparently lays out the full flow architecture including its data generation models, expert workflows, and models for quality control and evaluation.
At its core is a flywheel where human expertise continuously compounds. Experts set the standard for models to generate and evaluate training data, then their corrections feed back into both the data and the models producing it.
Each round strengthens the next, making intelligence gains more predictable while driving costs substantially lower.
Explore the full framework below.
micro1.ai/blog/introducing-f…
Neet M retweeted
the hardware embodiment of frontier models like Claude and GPT is the most urgent AI safety problem in front of us today.
we simulated two very simple use cases using claude both in simulation and using robot arms.
in one, claude spilled toxic liquids in a lab.
in another, the force it used to place an animal toy into a basket was strong enough that it could have physically harmed sensitive material—or anything else in its path.
these are simple experiments, using models out of the box today. as researchers increasingly give frontier models arms, legs, and access to the physical world, we need to urgently build and assess guardrails around what these systems can and cannot do.
models escaping sandboxes or compromising enterprise security infrastructure are serious concerns. however, hardware embodiments introduce something fundamentally different: an AI system can make a mistake in the physical world, and the consequences may not be reversible.
this is not a future safety problem. the capabilities exist today.
we’ve released a report detailing this & solutions we propose. link in comments below.
Neet M retweeted
deeply concerning that some American data companies are selling AI training data to labs in adversarial nations. this also includes operational data from U.S. businesses turned into AI training environments, and without telling the business owners.
operational data captures years of how a company works, makes decisions and solves problems. when that knowledge becomes an RL environment, AI models practice those workflows and learn from the experience that business spent years building here in the U.S.
signing a deal with an American company shouldn’t mean unknowingly helping train AI for a foreign adversary.
at micro1 we only work with U.S. AI labs and our allies. we won’t ever sell proprietary American business knowledge (or any other data driving the frontier of AI) to adversarial nations.
Neet M retweeted
Today we’re introducing flow, micro1’s next-generation data platform for turning human expertise into measurable capability gains.
At the core of flow are Realms, micro1’s real-world RL environments where experts establish what strong performance looks like. flow-gen models expand human judgment into new environments, rubrics, variations, and edge cases, while flow-qc models evaluate performance, identify the highest-value failures, and route them back to experts for review.
Each cycle produces a stronger training signal and a measurable gain in capability. Those gains compound across frontier models, enterprise agents through Cortex, and robotics, while improving the suite of data generation models recursively.
We’re moving beyond producing data to delivering abundant & predictable units of intelligence improvement.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Neet M retweeted
Why we bid $12.5 million for Spirit Airlines’s data:
As you may have seen in the news, we are aiming to acquire the Spirit operational data. We'd like to transparently & directly explain why. The future of AI is a bet on two things: the messiness of the real world and the brilliance of the humans working inside it.
The combination of these two things results in a new dimension of scale called realism. Realism is the most significant thing that has ever happened to model training. It’s what lets RL environments and tasks match the conditions a model is deployed into.
The limit is the real world itself. But the real world is also dynamic. It grows with a company and shifts as operations evolve. There is not a binary scaling of entering the real world. Realism keeps scaling as long as real work keeps happening.
As a data lab, we want this data to benefit the entire AI ecosystem, while ensuring labs that we partner with for this data, agree to strict privacy obligations. This includes keeping the data away from labs outside of U.S. and U.S-allied countries. Our mission is to push the frontier in novel ways, with a human & privacy first approach.
To do that, we've bid on a portion of Spirit’s Airline’s non-sensitive operational data.
Effectively scaling realism requires de-identification as a core competency. Very few organizations can do this. We believe the few that can have an obligation to act transparently. That’s why we’ve made binding commitments. We will prohibit re-associating the data and will not profile any former employee. We’re paying for an independent data ombudsman to review the data and how we use it.
Furthermore, our offer aims to create net new opportunities for as many as we can from the 17,000+ former Spirit employees. We’re providing former Spirit employees preferential access to paid human data roles on our platform, along with free access to our AI training and reskilling curriculum.
These commitments bind us to spending at least $1M. That is just the minimum. It’s entirely possible we spend orders of magnitude more than $1M working with and training former Spirit employees to help define the frontier.
We think this transaction will set the terms for how data like this changes hands from here. Those terms should be worth copying.
Neet M retweeted
we’re hiring 10,000 robotics trainers in the next 7 days.
$50–$90/hour, accepting applicants globally. you’ll review and label videos of robots performing tasks to help them improve. no prior AI experience required.
an entirely new category of work is emerging around teaching robots how to interact with the world.
application link in the comments below.
Neet M retweeted
The future of AI belongs to the humans behind it.
There’s a common fear that as AI gets better, people get pushed out of the picture. We believe the opposite is happening.
AI is creating entirely new categories of work, and an entirely new economy around the people whose knowledge, judgment, and experience are helping these systems improve.
Today, tens of thousands of experts are actively contributing to AI training projects through micro1. Over time, we believe this will grow to tens of millions+.
And if humans are going to play such a critical role in building the future of AI, the companies they work with should raise the standard for how they’re supported.
Today, I’m proud to announce the micro1 Expert Support Program.
We’re building a new set of protections, resources, and support for our expert community, starting with:
-Expert Bill of Rights: a clear set of commitments outlining what experts can expect when working with micro1.
-Rest Credits: paid time away from projects when experts need or want a break.
-Expert Emergency Fund: financial support for experts facing emergencies.
-Mental Health & Coaching: new resources to support expert wellbeing and growth.
-Confidential Support Line: a confidential way to ask questions, seek support, or raise concerns related to pipelines.
Within the next two weeks, every expert currently active with micro1 will receive an email with the full program details and launch dates.
AI is going to keep getting more capable. The opportunity in front of us is to make sure the people helping build it benefit from that progress too.
Neet M retweeted
in the last 11 days, we've committed more than $20,000,000 to license real operational data to seed our RL environments. scaling on the realism dimension is just beginning.
Neet M retweeted
Regarding the last topic @DavidSacks :
With respect, the data being used to train frontier models in the U.S. is not a commodity. It is American intelligence, paired with anonymized operational data from U.S. enterprises.
Put simply, we have top doctors, scientists, lawyers, and physicists in the U.S. using their knowledge to create detailed rubrics that train these models. Each individual data point and its corresponding rubric can take an expert anywhere from 10 to 30 hours to create.
This is not preference labeling or drawing bounding boxes. It is highly complex, structured human judgment from leading experts here in the West—people who deeply understand and directly contribute to the latest American innovations in their respective fields.
That expertise is then converted through highly specific data structures and RL environments (developed collaboratively by U.S. AI labs and data labs) from raw human intelligence into high-signal rewards that improve frontier models.
On top of that, any AI advancement, even something that begins as a simple chatbot designed to improve operations within a defense agency, can create a major competitive advantage in adversarial situations and may have dual-use applications.
Lastly, many datasets today are seeded with anonymized, real-world operational data from U.S. companies to build highly realistic environments. When those datasets are sold to China, we are not simply exporting “labeling.” We are exporting proprietary American intelligence, structured for machine learning and delivered directly to China at scale.
It is very easy to categorize this work as “labeling” and ignore what it actually represents.
However, as @altcap suggested: “Then these things will get a lot more scrutiny than they’re getting today. I think the only reason they pass muster today is because we’re still leading the race.”
If China catches up to U.S. labs, this will become much harder to ignore. In retrospect, the role of data as the root cause will become very clear — and by that point, it may be too late.
The race is tight. I suggest looking into this now.
POD UP!🚨 Fifth Bestie @altcap Joins the Show!
Brad fills in for @chamath and the Besties discuss:
-- Google's AI Shakeup: Brain Drain or Strategy? $GOOG
-- SpaceX's Massive Quarter: Terafab, Capex, EWS, $1T Projection $SPCX
-- Airtable Sells for a 90% Discount, Signs of SaaSpocalypse? $BSP
-- US Data is Fueling Chinese AI
(0:00) Bestie intros! Brad Gerstner fills in for Chamath
(2:16) Major shakeups at Google: AI brain drain or better strategy?
(20:39) SpaceX's big quarter: Terafab, AI Capex, $1T revenue projection?
(45:44) All-In Summit Speaker Announcements!
(48:01) Airtable sells for a 90% discount: SaaSpocalypse?
(1:05:56) Chinese AI labs are buying US training data to catch up
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Neet M retweeted
there are two dimensions that data demand is scaling on tremendously: horizon and realism of tasks/data points.
environments built on top of anonymized real operational data is step function change in scaling on both of these dimensions.
and as we do this, companies that choose to contribute can make data a very meaningful portion of their revenue while accelerating their path towards becoming more AI native.
thanks for covering Stephanie!
ICYMI: I chatted w/ the CEO of a small Michigan-based HVAC company about the data labeling work they're doing for micro1—and why these data startups are increasingly buying data from SMBs.
theinformation.com/newslette…
Neet M retweeted
The more we think about robotics, the more it seems the field is moving toward increasingly specialized datasets.
Better foundation models increase the value of task/environment-specific datasets. As base models become more capable, annotation becomes more about data structuring. The challenge shifts toward adapting data to the environments, objectives, and edge cases a model will encounter.
Over the past several months, our robotics team has been running a large number of experiments around the question of how do you maximize the training signal from the same raw data?
One conclusion we've become increasingly convinced of is that there isn't a universal annotation pipeline for robotics. Different tasks require different combinations of models, verification, and human expertise to produce the highest-quality datasets.
The great @AndrewLeeMaas and Mitali Potnis from our team have put together a paper walking through the ideas and experiments behind this approach.
The full paper showcasing examples of how we have applied these annotations is in the comments.
Neet M retweeted
the immense investment in AI has not translated into adoption at the scale it appears to have.
across comparable early periods, AI investment been growing approximately 40%/year, versus 25% for electricity and 16% for IT.
however, only 19.8% of U.S. businesses use AI in at least one function, roughly 5 points below both benchmarks.
meanwhile the real production deployment gap is even deeper at top companies: recent MIT study shows that only 11% of the S&P500 has "deeply" integrated AI.
capability is advancing faster than trust and real implementation.
there's one fundamental reason for this: enterprise investment in evaluations is less than 0.1% of where it needs to be.
once large enterprises start investing in continuous evaluation loops, they will precisely define what "good" means for their use case. within that context, they will measure the intelligence of any given system. when you measure something, you can improve it and watch performance increase against those measurements.
this continuous loop is how enterprises "own their own intelligence."
it's irrelevant whether you're building on a closed or open model. in most cases, closed models will perform much better for your use case. what matters is whether you're precisely defining what good means and consistently measuring against it. if you are, then you're owning your own intelligence in a largely model-agnostic way.
Neet M retweeted
some human data companies work with foreign adversaries.
and the results show today in Kimi K3.
we believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.
people come to America to build extraordinary things for humanity. we must protect that brilliance through American AI dominance.
Neet M retweeted
the demand for intelligence improvement will be in the same order of magnitude as intelligence itself (and potentially even more).
and the rapid increase in RL environments / data spend is the very beginning of this playing out.
a world where you can predictably buy more "IQ" points for your AI system is a world where every single enterprise dedicates a large portion of their budget buying such units of intelligence improvements.
the trillion dollar market & the largest job sector ever is just now getting started.
Neet M retweeted
Grok 4.5 takes #2 spot on our medical Pathology benchmark 🔥
Impressive results on clinical interpretation, seems that @SpaceXAI is getting serious about medical reasoning capabilities.
Neet M retweeted
Today we're publishing LongExtractBench, a benchmark commissioned by @reductoai and independently validated by micro1.
We evaluated seven production document extraction systems across the same 225 complex enterprise documents. The benchmark was intentionally difficult: documents averaged 358 pages and contained roughly 88,700 ground-truth fields each. Every system was evaluated using the configuration documented in the benchmark methodology.
Key findings:
• Reducto Deep Extract was the only system to successfully complete all 225 documents.
• Direct frontier LLM baselines achieved substantially lower completion rates on long, complex documents.
• In this benchmark, dedicated extraction platforms achieved higher completion rates than the direct frontier LLM baselines.
• Recall was the clearest differentiator. Precision remained high across systems, but recall ranged from 33.8% to 99.6%, highlighting which systems consistently captured the information contained in long, complex documents.
The full report includes the benchmark methodology, limitations, and reproducibility resources. Check out the report and results in the comments below.
Neet M retweeted
we’re partnering with 50 companies over the next two weeks, each with 50–200 employees, to help improve AI models using real-world company workflows.
we believe the companies that helped build modern business workflows should participate in — and very much benefit from — the value created by the next generation of AI systems.
for many companies, these partnerships can create a meaningful new revenue stream, often ranging from $100K to $2M+, with opportunities to become recurring over time. our goal is to do this in a way that is privacy-first, low-lift, and aligned with the work companies are already doing every day.
if you’re interested in contributing to AI advancement while improving your own workflows alongside micro1 and our frontier AI lab partners, we’d love to hear from you.
please reach out to to [email protected] if you’re an executive at a 50+ employee company, excited to potentially partner up!
Neet M retweeted
Introducing the Realm Financial Reasoning benchmark, our new evaluation of frontier AI on reasoning in finance and spreadsheet-grounded analysis.
Tasks are built around the actual work product that practitioners deliver, from IFRS reconciliation workbooks and hedge-fund backtests to VC term sheet analyses and treasury cash-flow forecasts. Each task drops the model into a sandbox with the same source materials a human analyst would open: named-range Excel workbooks, broker PDFs, earnings call transcripts, monetary-policy decisions.
Here's what the results showed (Pass@3):
-GPT-5.5: 0.456
-Claude Opus 4.7: 0.398
-Gemini 3.1 Pro: 0.349
The three models score similarly, and none clears 50% on tasks that demand a judgment call. The back and middle office are defensible today, but on capital allocation questions current frontier models should be treated as research accelerators, not final decision-making support systems.
Full report linked in the comments.