@TrainyAI

Building open source tools for distributed training.

Joined June 2023
Over the past five years, I can't count the amount of times I've heard this said about @ClickHouseDB. Now we're bringing that same focus on speed and reliability to Postgres, to create the unified data platform that AI needs.
.Love testimonials like this! Brings a ton of validation to the hard work team has been putting to build the world's fastest and most reliable Postgres service. Thank you @asaiUwU for the kind words and all the feedback during the migration. 🙏 @ClickHouseDB
2
13
1
41
4,508
Thank you @TrainyAI team for all the trust in migrating from RDS Postgres to ClickHouse Managed Postgres. Really appreciate the feedback through the process and glad that performance has been much better on both OLTP and OLAP fronts, post migration. @roanakb @asaiUwU @ClickHouseDB
2
8
501
.Love testimonials like this! Brings a ton of validation to the hard work team has been putting to build the world's fastest and most reliable Postgres service. Thank you @asaiUwU for the kind words and all the feedback during the migration. 🙏 @ClickHouseDB
10
2
28
7,170
Trainy retweeted
Huge thanks to @ClickHouseDB for the spotlight! We're building the fastest experiment tracker out there and it wouldn't be possible without them. One user's query took 10-15 min on W&B, 5 seconds on Pluto. Read below 👇
2
9
501
Trainy retweeted
NeptuneAI shuts down March 5th. @TrainyAI just launched Pluto on @ycombinator, a drop-in replacement so you don't lose years of experiment data. Swap one import. Dual-log to validate. Export your history. Open source. On Neptune's official transition hub. ycombinator.com/launches/PLM…
3
2
20
7,258
Trainy retweeted
@TrainyAI's Konduktor platform helps bring the benefits of a leading research team to your GPU cluster. We provide a fault-tolerant scheduler, integrated observability, and more. Check out our docs: konduktor.readthedocs.io/en/…
2
2
258
Trainy retweeted
This leads to significantly higher (>80%) GPU usage. Add in some fault-tolerance to the infrastructure, and we see: - No more manual restarts at 2am. - ML Engineers get to focus on their jobs, rather than becoming DevOps experts.
1
1
2
243
Trainy retweeted
Top tier AI research teams (Meta, OpenAI, etc.) have figured out the most efficient way to work with a cluster of GPUs. Instead of managing each GPU separately, they create a pools of GPU nodes and let sophisticated schedulers manage GPU availability efficiently.
1
2
2
246
Trainy retweeted
Is your team struggling with GPU failures? Let’s talk! Docs: konduktor.readthedocs.io/en/…
1
1
125
Trainy retweeted
At @TrainyAI, we built a controller within Konduktor to monitor GPU node health and isolate unhealthy nodes. This way if a job fails, 0 manual intervention is required. K8s does its magic of placing work only on healthy nodes, and we forward relevant GPU/NCCL logs to your CSP. 🚀
1
1
1
113
Trainy retweeted
ML engineers shouldn’t be wasting time debugging infrastructure — especially when H100s have a 25-30% fault rate. 🛠️ ML infrastructure should be able to handle bumps and bruises to the underlying hardware.
1
2
2
153
Trainy retweeted
4/ Struggling with multinode setups on your cloud provider? We'll cut your setup time from weeks to minutes. Docs: konduktor.readthedocs.io/en/…
1
1
69
Trainy retweeted
3/ One of the biggest value-adds of @TrainyAI's Konduktor platform is that we simplify this complexity. We abstract away network configurations, so you can launch multinode training with high-bandwidth networking across different clouds in the same way.
1
1
1
71
Trainy retweeted
2/ At @TrainyAI, we've seen AI research teams lose over $10,000 trying to scale out due to misconfigured GPU fabrics. That's a costly mistake that can be avoided.
1
1
1
45
Trainy retweeted
Setting up and validating GPU networking is a lot less trivial than you'd think. Here's why: 1/ GPU fabric technology varies a lot across cloud providers for the H100. For example, Google Cloud has TCP-X, while AWS uses EFA. Once you commit to one setup, it often locks you in.
1
2
2
159
Trainy retweeted
He lays out the ARC-AGI benchmark, how it tests generalization abilities rather than memorization, and his thoughts on what kind of AI system will be necessary to improve on the SoTA. Watch here: youtube.com/watch?v=s7_NlkBw…
1
3
108
Trainy retweeted
3. Skill does not show intelligence. And displaying skill at any number of tasks does not show intelligence. - This misguided view of intelligence is what causes our current form of benchmarking to be inadequate.
1
1
2
61
Trainy retweeted
2. For any LLM, for any query that seems to work, there exists an equivalent rephrasing of the query that will break. - This ties into LLM's inability to handle deviations from a pattern - Highlights the modern LLM's lack of robustness
1
1
2
40
Trainy retweeted
1. The core limitations of Transformer-based architectures have not changed in over 5 years. - Inability to adapt to small deviations from memorized patterns - Weak, patchy generalization
1
1
2
39