@itsnextwork

The best way to learn AI and build a portfolio. Link below to access the projects👇

Texas, USA
Joined July 2024
WHAT IS THE IKEA ASSEMBLY BENCHMARK? An independent AI research group, Epoch AI, built a new test to see whether AI models can spot mistakes in half-built furniture. They bought three IKEA products (a shoe cabinet, a bed frame, a dresser), assembled them, and photographed 60 points along the way, deliberately introducing a visible mistake in some photos (like a panel installed backwards). Each AI model gets a photo, the official IKEA instruction manual, and a zoom tool, then has to say whether the build is correct so far, and if not, exactly which step went wrong and why. WHY ARE PEOPLE TALKING ABOUT IT? Because the score jump is dramatic. In November 2025, the best model, Anthropic's Claude Opus 4.5, got only 28% right, barely above guessing. Ten months later, per Epoch AI's results published September 23, 2026, OpenAI's new GPT-6 Astra scored 80%, and did it faster too, about 3 minutes per photo versus far longer for every other model tested. Claude Fable 5.1 came in at 70%, Claude Opus 5 at 61%. Tech outlets picked it up fast this week (The Decoder ran it under the headline "GPT-6 Astra can now tell you exactly where you screwed up your IKEA shelf"), and it's been spreading through AI newsletters since. WHAT ARE PEOPLE BUILDING, POSTING, OR DOING WITH IT? This is a research benchmark, not a shipped product, so right now the activity is researchers and AI writers picking apart the results rather than people using it directly. Epoch AI published a full leaderboard and flagged a clear failure pattern: Google's Gemini and Alibaba's Qwen tend to falsely flag correct builds as wrong, while earlier GPT and Claude models tended to miss real mistakes outright. Commentators are treating this as a case study in how fast "spatial reasoning" (matching a photo to a diagram) is improving, and pointing to car repair and appliance troubleshooting as the next real-world versions of the same skill. HOW COULD IT AFFECT YOU? Not directly, not yet. You can't point your phone at your own bookshelf and get an answer today; this benchmark covers just 3 specific IKEA products and lives inside a research report, not a consumer app. It's a preview of a skill, comparing a photo to instructions and finding the difference, that could eventually turn up inside apps meant to help you fix things around the house. For now, this is happening behind the scenes, inside AI labs and evaluation reports, not in something you'd open yourself.
2
3
72
Why you can't find a job ❌ Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
1
37
9 MCP servers you should install this week ⚡ Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
1
55
5 Claude projects you can finish this weekend 👀 Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
2
49
Software engineer: $136K. Web developer: $85K. What creates the $50K gap? Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
1
2
52
WHAT IS JEV? Jev is a new AI model from a startup called TypeSafe AI, released September 15, 2026. Unlike ChatGPT-style chatbots, Jev doesn't write sentences. You give it a short piece of text, like a support ticket or an email, plus a specific yes/no or multiple-choice question, and it hands back a typed answer with a confidence score, what the company calls a "calibrated decision." Example: feed it the sentence "The support agent issued a full refund to the customer" and ask "Was a refund issued?" It returns something like true, 97% confident, instead of writing out a paragraph. WHY ARE PEOPLE TALKING ABOUT IT? Jev was built by Diogo Almeida, who spent about four years at OpenAI helping build ChatGPT and co-inventing RLHF (reinforcement learning from human feedback), the training method behind most chatbots today. He has said models trained to please people in conversation are clumsy at the narrower, structured decisions software actually needs to run. TypeSafe launched Jev alongside a $40 million seed round led by DCVC, which Forbes reported valued the company at $200 million. According to Vercel's own blog post, Jev reached about 13% of Vercel's paid teams within 24 hours of launching on its AI Gateway, about twice the adoption of OpenAI's GPT-5.6 family and more than six times Fable 5.1 in that same window. WHAT ARE PEOPLE BUILDING, POSTING, OR DOING WITH IT? Vercel CEO Guillermo Rauch posted that after swapping OpenAI's Luna model for Jev in a system that checks whether commands are safe to run, results came back up to 18 times faster and more accurate, by his own account. Bryo AI's CTO, Nikhil Mudholkar, ran a side-by-side test classifying business emails: Google's Gemini was slightly more accurate but 10 to 20 times more expensive, and what stood out to him was that Jev was the only one that returned a real probability score instead of just a label. Other developers are testing Jev as a cheap "sanity checker" that watches over pricier chatbot-style agents and flags when they might be going off track. HOW COULD IT AFFECT YOU? This is developer infrastructure, not something you'd open an app to use directly. But you may experience it behind the scenes: when a company's software routes your support request to the right team, decides whether your message needs a human, or double-checks that an AI agent's action is safe before it happens, tools like Jev could be the quiet, cheap layer making that specific yes/no call, while a slower chatbot model still handles the parts meant to talk to you.
1
1
134
The tech interview secret 🤫 Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
41
What a hiring manager sees when every resume used AI 👀 Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
1
41
Data Analyst vs Data Scientist: what's the difference? 📈🧬 Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
29
The 7 levels of system design 🏗️ Follow and comment "projects" for hands-on projects to add to your resume (step-by-step guide included)! #AI #TechTok #TechCareers #CareerTips #FutureOfWork nextwork.ai/?utm_source=x&ut…
1
43