@PhilForrencei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Hello
Philadelphia, PA
Joined July 2023
- Tweets6.1K
- Following191
- Followers225
- Likes9.1K
Phillip Forrence retweeted
OpenAl is scrapping the release of its next-generation Al model (GPT-6.1 Astra) over safety concerns that researchers raised during internal testing. wsj.com/tech/ai/openai-chatg…
Phillip Forrence retweeted
Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen
With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2 on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task)
Key takeaways:
➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it
➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max)
➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task
➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus
These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon.
Other model details:
➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5
➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2
➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
Phillip Forrence retweeted
Sonnet 5.5 is smarter, more efficient, and 30% faster than Sonnet 5. It costs up to 30% less for most work, so your Claude Code usage goes further too.
Use it for well-scoped everyday tasks like fixing bugs and quickly iterating on features.
Phillip Forrence retweeted
Today we launched Claude Sonnet 5.5! 30% faster and costs up to 30% less than Sonnet 5 for most work.
Strong for well-scoped everyday tasks (fixing bugs, creating docs, slides). @RLanceMartin had models paint the same photo in code. You can see the jump:
Phillip Forrence retweeted
September 28, 2008: SpaceX's Falcon 1 rocket reaches orbit for the first time.
September 28, 2026: SpaceX's Starship rocket reaches orbit for the first time.
someone's going viral after copying my exact style.
honestly that's just proof the design works.
I bet he saw this post and asked claude to replicate it. it's flattering :)
I'm going to share exactly how to do it anyway,
so anyone can build 3d interactive STEM lessons.
pls try it with different subjects.
make it useful for anyone who wants to learn.
one thing I noticed: his page took a while to load
because it was stuffed with too much.
my goal is simple:
it needs to load really fast.
it should be lightweight.
it should be useful for learning.
it should be visually stunning.
AI is cool, but these 3d interactive lessons can still produce inaccurate results right now,
so they really need input from experts.
an open-source repo is coming!
I'm hoping people will add new learning material or fix inaccurate visuals just by prompting.
it needs to be done properly.
let me know what you want to see??
spoiler: it's never about crafting the perfect prompt or /skills.
you won't need any of that.
where we're going, you just ask for stuff.
I asked Opus 5.5 to explain camera focus by building an interactive lens lab
Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost
lens.lab.sael.net
Move the focus ring and you can see the glass elements shift the sharp plane through the scene
Phillip Forrence retweeted
Crazy stat of the day: Each one of these V3 Starlink satellites being deployed right now increases the bandwidth capacity of the ENTIRE Starlink constellation by roughly 0.11%.
That’s one satellite.
For perspective, there are ~11,100 Starlink satellites in orbit today, mostly V2 Minis. Now imagine hundreds of Starship launches carrying 60 V3 satellites each. That’s why @elonmusk says that it's increasingly probable that Starlink will carry a majority of Earth’s IP traffic long-term.
SpaceX's Starship rocket has successfully reached orbit and deployed its V3 Starlink satellites into orbit for the first time, officially making this Starship's first revenue-generating flight, a major milestone for the company!
Congrats @SpaceX team! Incredible achievement 🚀
Phillip Forrence retweeted
SpaceX's Starship rocket just performed a flip maneuver and splashed down on target in the Pacific Ocean. Seems like tower catch could be a go for flight 15.
Phillip Forrence retweeted
With today’s single Starship launch, SpaceX added more Starlink capacity in the last hour than the first ~1.5 years of Starlink launches. The next launch will cover another full year. Welcome to the network, V3! 🚀
Phillip Forrence retweeted
Claude Code Opus5.5にリアルタイム海岸シミュレーションを作って!!と言った結果
カメラに水滴がついたり、海の中に魚が居たり、鳥が飛んでたりする👀
UltracodeとMax
Three.js使用
#claude
Phillip Forrence retweeted
Not sure what happened to Claude limits they went from barely usable to basically unlimited 😂
Phillip Forrence retweeted
Glad Claude limits feel different!
The lower Opus 5.5 price goes straight into your limits. They go about 25% further than Opus 5. Cache reads cost 60% less and output is generated 30% faster:
claude.dev/blog/what-a-task-…