@Nipsulii
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Product Engineer | ML & AI Systems | Founder Co-Founder at 041
Baltimore, MD
Joined July 2009
- Tweets877
- Following820
- Followers111
- Likes5.3K
Thank @jullerino for great example for my agent to find on how to do Clerk device flow
What a time to be alive! My agent optimized ML framework just got on par or passed PyTorch on our performance benchmark and another agent using the framework got to spitting distance of our reference benchmark with our new architecture.
Seems that one can solve problems by just throwing tokens at them
I've been building new ML training framework because I got annoyed about all the the mistakes the agents were doing when implementing training scripts and models in python.
They didn't put the focus where they should have. To keep us moving the experiments forward I've been building really strict "research factory" system that contains scripts, and issue templates and review rules and all kinds of stuff to try to force the models to behave certain way.
And today I realized that if my new lisp based ML framework lives to it's promise it's going to make most of those things obsolete! I'm so excited to start the real world experiments with the new system and do evals on how well it actually helps in making ML research at the speed of thought.
As a technical founder and a dad the only thing that allows me to be as productive I am is @t3dotcodes mobile app.
Does your ML experiment tracking tool do this?
I've spend few months trying to figure out how to do AI assisted AI research the most efficient way. Been building tooling to help us move faster. I've tried different things to provide different interfaces for the agents to explore and diagnose the training runs.
I build cli, 2 versions of API, python SDK for easy querying, was thinking about MCP. Nothing seemed to feel right. Outside of loss curves the diagnostics were always case dependent analysis. Then it hit me. Why don't I just expose the collected metrics for the agent via SQL?
No more random data sync scripts, no more random cli command pipes created by the agents to find the insights on why does the gradient vanish at layer 7. Just raw SQL that the agents know. Accessed data cached locally, series data in parquet so one can query it also outside of duckdb.
I built a custom ML metric tracking system for our AI lab. The rational was to optimize it for the research agents, it has cli, nice api, nice python SDK.
I rarely look at the metics on the dash, I always have agent to write script and custom visualization. Often syncing data to local duckdb to speed things up.
It started to feel wasteful so decided to write @duckdb extension for it to speed up the agentic AI research even more. Just throw tokens at the problem like @theo suggests.
I don't know why I'm really doing this. But building new ML training framework from scratch with lisp frontend. This is the 4th from scratch (or close from scratch) iteration that I'm having on this thing. Now with clear two phase compiler with well defined IR in between.
The idea came from seeing the agents doing stupid mistakes again and again with the python training stuff I thought what if I created DSL + clear convention based experiment setup. But something that would still allow crazy expressivity of ideas.
And the more I thought about it it always came back to lisp. All the 4 iterations have had list frontend and some sort of native backend, but those have ended up being massive slop fests. Let's see if the 4th iteration is the one that works well.
Have done few threads on @t3dotcodes damn this is massive productivity booster. Mainly the remote capabilities. So far 2 old macs seems to have been enough for remote work.
The blog hits so hard. I had left codex on a loop to solve a problem but the progress seems to be nowhere and then asked fable to check the work and it turned out to be almost 400k lines of irrelevant slop.
Astra is really, really cool but I cannot current trust it for my present day engineering. At least until I have adjusted. Some random thoughts including code samples of what I recovered from my traces. lucumr.pocoo.org/2026/9/7/as…