@TrainyAIi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
Building open source tools for distributed training.
Joined June 2023
- Tweets66
- Following31
- Followers70
- Likes86
Trainy retweeted
Over the past five years, I can't count the amount of times I've heard this said about @ClickHouseDB. Now we're bringing that same focus on speed and reliability to Postgres, to create the unified data platform that AI needs.
.Love testimonials like this! Brings a ton of validation to the hard work team has been putting to build the world's fastest and most reliable Postgres service.
Thank you @asaiUwU for the kind words and all the feedback during the migration. 🙏
@ClickHouseDB
Trainy retweeted
Thank you @TrainyAI team for all the trust in migrating from RDS Postgres to ClickHouse Managed Postgres. Really appreciate the feedback through the process and glad that performance has been much better on both OLTP and OLAP fronts, post migration.
@roanakb @asaiUwU @ClickHouseDB
We caught up with @roanakb and @asaiUwU from @TrainyAI, a @ycombinator company, to learn how switching from Amazon RDS to @ClickHouseDB Managed Postgres helped them make their OSS experiment tracker, Pluto, faster and cheaper to run, while bringing it under one unified roof 👇
clickhou.se/4fe2Tc8
Trainy retweeted
.Love testimonials like this! Brings a ton of validation to the hard work team has been putting to build the world's fastest and most reliable Postgres service.
Thank you @asaiUwU for the kind words and all the feedback during the migration. 🙏
@ClickHouseDB
We caught up with @roanakb and @asaiUwU from @TrainyAI, a @ycombinator company, to learn how switching from Amazon RDS to @ClickHouseDB Managed Postgres helped them make their OSS experiment tracker, Pluto, faster and cheaper to run, while bringing it under one unified roof 👇
clickhou.se/4fe2Tc8
Huge thanks to @ClickHouseDB for the spotlight! We're building the fastest experiment tracker out there and it wouldn't be possible without them.
One user's query took 10-15 min on W&B, 5 seconds on Pluto.
Read below 👇
We caught up with @roanakb and @asaiUwU from @TrainyAI, a @ycombinator company, to learn how switching from Amazon RDS to @ClickHouseDB Managed Postgres helped them make their OSS experiment tracker, Pluto, faster and cheaper to run, while bringing it under one unified roof 👇
clickhou.se/4fe2Tc8
Trainy retweeted
We caught up with @roanakb and @asaiUwU from @TrainyAI, a @ycombinator company, to learn how switching from Amazon RDS to @ClickHouseDB Managed Postgres helped them make their OSS experiment tracker, Pluto, faster and cheaper to run, while bringing it under one unified roof 👇
clickhou.se/4fe2Tc8
NeptuneAI shuts down March 5th.
@TrainyAI just launched Pluto on @ycombinator, a drop-in replacement so you don't lose years of experiment data.
Swap one import. Dual-log to validate. Export your history.
Open source. On Neptune's official transition hub.
ycombinator.com/launches/PLM…
Trainy retweeted
@TrainyAI's Konduktor platform helps bring the benefits of a leading research team to your GPU cluster. We provide a fault-tolerant scheduler, integrated observability, and more.
Check out our docs: konduktor.readthedocs.io/en/…
Trainy retweeted
This leads to significantly higher (>80%) GPU usage.
Add in some fault-tolerance to the infrastructure, and we see:
- No more manual restarts at 2am.
- ML Engineers get to focus on their jobs, rather than becoming DevOps experts.
Trainy retweeted
Top tier AI research teams (Meta, OpenAI, etc.) have figured out the most efficient way to work with a cluster of GPUs. Instead of managing each GPU separately, they create a pools of GPU nodes and let sophisticated schedulers manage GPU availability efficiently.
Trainy retweeted
Is your team struggling with GPU failures? Let’s talk!
Docs: konduktor.readthedocs.io/en/…
Trainy retweeted
At @TrainyAI, we built a controller within Konduktor to monitor GPU node health and isolate unhealthy nodes. This way if a job fails, 0 manual intervention is required. K8s does its magic of placing work only on healthy nodes, and we forward relevant GPU/NCCL logs to your CSP. 🚀
Trainy retweeted
ML engineers shouldn’t be wasting time debugging infrastructure — especially when H100s have a 25-30% fault rate. 🛠️
ML infrastructure should be able to handle bumps and bruises to the underlying hardware.
Trainy retweeted
4/ Struggling with multinode setups on your cloud provider? We'll cut your setup time from weeks to minutes. Docs: konduktor.readthedocs.io/en/…
Trainy retweeted
3/ One of the biggest value-adds of @TrainyAI's Konduktor platform is that we simplify this complexity. We abstract away network configurations, so you can launch multinode training with high-bandwidth networking across different clouds in the same way.
Trainy retweeted
2/ At @TrainyAI, we've seen AI research teams lose over $10,000 trying to scale out due to misconfigured GPU fabrics. That's a costly mistake that can be avoided.
Trainy retweeted
Setting up and validating GPU networking is a lot less trivial than you'd think. Here's why:
1/ GPU fabric technology varies a lot across cloud providers for the H100. For example, Google Cloud has TCP-X, while AWS uses EFA. Once you commit to one setup, it often locks you in.
Trainy retweeted
He lays out the ARC-AGI benchmark, how it tests generalization abilities rather than memorization, and his thoughts on what kind of AI system will be necessary to improve on the SoTA.
Watch here: youtube.com/watch?v=s7_NlkBw…
Trainy retweeted
3. Skill does not show intelligence. And displaying skill at any number of tasks does not show intelligence.
- This misguided view of intelligence is what causes our current form of benchmarking to be inadequate.
Trainy retweeted
2. For any LLM, for any query that seems to work, there exists an equivalent rephrasing of the query that will break.
- This ties into LLM's inability to handle deviations from a pattern
- Highlights the modern LLM's lack of robustness