iAccount based inViet Nam!
About this account
- Account based in
- Viet Nam
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
The AI OS for scientific discovery. We run the whole research lifecycle — ideas, experiments, IP, funding, publication. Built for world-class science.
- Tweets42
- Following22
- Followers3.5K
- Likes17
ALT 72.24 minutes to GPT-2 on 8×H100. AutoTrust's ScienceGuru, running Guru Turbo 1.2, posted the fastest result we've found on Andrej Karpathy's Time-to-GPT-2 benchmark (one run, self-reported): 27% under the official record and 9.6 min ahead of the best community recipe. The field Reported times on the official leaderboard and in public nanochat PRs, as of Sept 24: - 72.24 min · ScienceGuru (AutoTrust) · 1 run - 81.84 min · Giovanni Zinzi · 6 runs - 91.74 min · Oriole Networks · 3 runs - 94.58 min · Martin Jurča (Seznam.cz) · 6 runs - 94.6 min · Weco-optimized run · 3 runs - ~99 min · Official record, Karpathy's autoresearch round 2 · 5 runs Zinzi also reported a 73.92-min experiment with a Cosmopedia data mix (3 runs), which he set aside after it regressed at a smaller model size. The lead - 27% faster than the official record (−26.8 min) - 12% faster than the best community recipe (−9.6 min) - 19–22 min ahead of Oriole Networks, Seznam.cz and the Weco-optimized run How Starting fr
ALT 72.24 minutes to GPT-2 on 8×H100. AutoTrust's ScienceGuru, running Guru Turbo 1.2, posted the fastest result we've found on Andrej Karpathy's Time-to-GPT-2 benchmark (one run, self-reported): 27% under the official record and 9.6 min ahead of the best community recipe. The field Reported times on the official leaderboard and in public nanochat PRs, as of Sept 24: - 72.24 min · ScienceGuru (AutoTrust) · 1 run - 81.84 min · Giovanni Zinzi · 6 runs - 91.74 min · Oriole Networks · 3 runs - 94.58 min · Martin Jurča (Seznam.cz) · 6 runs - 94.6 min · Weco-optimized run · 3 runs - ~99 min · Official record, Karpathy's autoresearch round 2 · 5 runs Zinzi also reported a 73.92-min experiment with a Cosmopedia data mix (3 runs), which he set aside after it regressed at a smaller model size. The lead - 27% faster than the official record (−26.8 min) - 12% faster than the best community recipe (−9.6 min) - 19–22 min ahead of Oriole Networks, Seznam.cz and the Weco-optimized run How Starting fr
ALT 72.24 minutes to GPT-2 on 8×H100. AutoTrust's ScienceGuru, running Guru Turbo 1.2, posted the fastest result we've found on Andrej Karpathy's Time-to-GPT-2 benchmark (one run, self-reported): 27% under the official record and 9.6 min ahead of the best community recipe. The field Reported times on the official leaderboard and in public nanochat PRs, as of Sept 24: - 72.24 min · ScienceGuru (AutoTrust) · 1 run - 81.84 min · Giovanni Zinzi · 6 runs - 91.74 min · Oriole Networks · 3 runs - 94.58 min · Martin Jurča (Seznam.cz) · 6 runs - 94.6 min · Weco-optimized run · 3 runs - ~99 min · Official record, Karpathy's autoresearch round 2 · 5 runs Zinzi also reported a 73.92-min experiment with a Cosmopedia data mix (3 runs), which he set aside after it regressed at a smaller model size. The lead - 27% faster than the official record (−26.8 min) - 12% faster than the best community recipe (−9.6 min) - 19–22 min ahead of Oriole Networks, Seznam.cz and the Weco-optimized run How Starting fr
ALT What happens when an AI research system picks up where two years of human optimization left off? The benchmark is the NanoGPT Speedrun: train GPT-2 to 3.28 validation loss on FineWeb, the target set by @karpathy 's llm.c replication, which took 45 minutes to get there. The speedrun's code descends from llm.c's PyTorch trainer, itself descended from NanoGPT, hence the name. Over two years, 91 official records brought the time down to 67.56s. ScienceGuru just hit 24.90s on 8×H100 (five seeds, self-reported): 108× faster than where the benchmark started, 2.71× faster than the official record, and 3× faster than Recursive's June record of 75.4s. Earlier this month it also took #1 on Autoresearch@Home at 0.8895 BPB, ahead of @Recursive_SI 's 0.9109, and the validated lead on @MedARC_AI 's NanoPath v2. We built on open PRs by Deven Pietrzak (ANVIL2) and Herman Brunborg (Exact-match), credited in the repo. ScienceGuru's RSI recipe fused them, cut the schedule from 1,194 to 652 steps and r
ALT What happens when an AI research system picks up where two years of human optimization left off? The benchmark is the NanoGPT Speedrun: train GPT-2 to 3.28 validation loss on FineWeb, the target set by @karpathy 's llm.c replication, which took 45 minutes to get there. The speedrun's code descends from llm.c's PyTorch trainer, itself descended from NanoGPT, hence the name. Over two years, 91 official records brought the time down to 67.56s. ScienceGuru just hit 24.90s on 8×H100 (five seeds, self-reported): 108× faster than where the benchmark started, 2.71× faster than the official record, and 3× faster than Recursive's June record of 75.4s. Earlier this month it also took #1 on Autoresearch@Home at 0.8895 BPB, ahead of @Recursive_SI 's 0.9109, and the validated lead on @MedARC_AI 's NanoPath v2. We built on open PRs by Deven Pietrzak (ANVIL2) and Herman Brunborg (Exact-match), credited in the repo. ScienceGuru's RSI recipe fused them, cut the schedule from 1,194 to 652 steps and r
ALT What happens when an AI research system picks up where two years of human optimization left off? The benchmark is the NanoGPT Speedrun: train GPT-2 to 3.28 validation loss on FineWeb, the target set by @karpathy 's llm.c replication, which took 45 minutes to get there. The speedrun's code descends from llm.c's PyTorch trainer, itself descended from NanoGPT, hence the name. Over two years, 91 official records brought the time down to 67.56s. ScienceGuru just hit 24.90s on 8×H100 (five seeds, self-reported): 108× faster than where the benchmark started, 2.71× faster than the official record, and 3× faster than Recursive's June record of 75.4s. Earlier this month it also took #1 on Autoresearch@Home at 0.8895 BPB, ahead of @Recursive_SI 's 0.9109, and the validated lead on @MedARC_AI 's NanoPath v2. We built on open PRs by Deven Pietrzak (ANVIL2) and Herman Brunborg (Exact-match), credited in the repo. ScienceGuru's RSI recipe fused them, cut the schedule from 1,194 to 652 steps and r
ALT Turbo 1.2 Preview is now live on ScienceGuru Web. ⚡ Guru Plan subscribers can try it free from September 15, 17:01 to September 19, 17:00 Singapore Time. From September 19, 17:01 to October 1, 17:00 Singapore Time, Turbo 1.2 Preview usage is 50% off. Log in, select Turbo 1.2 Preview from the model picker, and put it to work. https://scienceguru.ai/