@easemizei
iAccount based inBangladesh!
About this account
- Account based in
- Bangladesh
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
I joined the Dicky List. No thoughts, just Dicky. @DickyNFTs nitter.cf/DickyNFTs/status/21027…
yeah officer so he was using astra, and i told him he should've used fable, because fable writes more mergeable code, you know how gpt models love their tests? i also saw him on the terminal, we have guis now, officer have you heard about t3ode? it's the all in one interface for using your codex and claude subscription
Published 2 landing page templates on @21st_dev today.
Both made a sale on the first day.
Thanks 21st.
Templates in thread.
contra.com/products/kOgRgKDs… (On Discount for a limited time. price jumps double after 20th sep)
contra.com/products/ePM3HJD3…
Proof that Material Design is not bad. It's just in the bad hands
nitter.cf/lnkiai/status/20991264…
M3E Canvasが⭐️6k+ starsを突破🎉
ありがとうございます!
UIを全面的に見直し改善中です!大幅アップデートを楽しみにお待ちください!
github.com/lnkiai/m3e-canvas
Am I dreaming
How is SamA and D(m)ario on the same page. Something fishy or ...
Colors and typography are 40% of the design.
Don't let an LLM handle that part. leave it to an expert or
A Perfected algorithm
Introducing IMG2M3
No AI just a devtool that fixes the part AI can't
link👇
img2m3.vercel.app
AI's are good at literally everything right now but theres still this one part they can never nail.
Choosing the right colors for your apps.
Yeah they can choose maybe a one or two colors but what about the 32+ color mapping across dark and light theme your app is using.
This is where IMG2M3 comes.
It utilizes the #MaterialDesign Dynamic theme Generation algorithm to generate 5 modern themes across light modes. and maps that to #shadcn Color tokens.
No AI guessing - No looks good to me
Pure Algorithmic Perfection
You might like nitter.cf/winglee/status/2097750…
If You've used
Paseo (Paseo.sh)
Synarara (trysynara.com)
T3 Code (t3.codes)
Orca ADE (onorca.dev)
and 10+ other
Openai code red might happen again when Gemini 4 releases on 2030
Yes Artificial analysis literally changed the whole evaluation model to rank astra higher than MUSE
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
Intelligence Index v4.2 changelog:
+ AA-Briefcase, our agentic knowledge work evaluation with a private test set
+ @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages
- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated
… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness
This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January.
We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.
Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!
Intelligence Index v4.2 changes in detail:
➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.
➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.
➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.
Key results:
➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google
➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier
➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
I guess a new GPUI component library trend will soon start
Replying to @huacnlee
Hey Jason, we will start working on infra to automatically publish GPUI releases by the end of the month. For now, GPUI will remain in the Zed repo, but we're aiming for more frequent releases as a goal.