@Ttp7Ei
iAccount based inJapan
About this account
- Account based in
- Japan
- Connected via
- Japan App Store
Account-level information from X, not a live location or the device used for a specific post.
数学をやっている人です。よろしくどうぞ(*.ˬ.)" AIで、イラストも作ってます✨
@grok
Joined December 2025
- Tweets1.2K
- Following2.1K
- Followers128
- Likes2.4K
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
Introducing a new generative model, Empirical Variational Autoencoder (EVA):
We can revive VAEs for sequential generation only with an extra single linear layer that predicts the next prior without VQVAEs or diffusion models.
Project page:
mapooon.github.io/EVAPage/
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
ClaudeCodeやCodexで使える表現のライブラリ「NANKA DEKIRU」を作りました
なんかできる映像表現やモーション、編集などをまとめてます
使いたい表現を選んでお使いのエージェントにコピペする機能つき!
公式サイトの「公開ツール」にあります!
esoragototv.com/
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
ちなみに意外と知られてないんですけど,
オーケストレーターモデルさえ無検閲なら別にサブエージェントモデルは普通のOSSだけでいいので,
ギリギリ10~20toks(まあ30は欲しいけど)でさえGLM-5.3 Uncen動かせれば,
あとはサブエージェントはAPIで爆安爆速のDeepSeek V4.1でもGLM-5.3-Flashでも(遅いけど)MiMo V2.6でも使ってスケールできるんすよ
なので,100万でも積めば脆弱性ハント工場作って初月ぶん回すなんて簡単にできてしまう
で,初月動かせばもう100万回収できるからさらにスケールできてしまう
無検閲モデルはrefusalの癖を自分で分析して指示の分割方法や内容を自分で工夫できるので,そこをうまくやるようにPiかなんかベースでハーネス組めばよろしい
もはや無理なんすよ止めるのは
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
New research 🚀 from the ACE (Agentic Context Engineering) team at Stanford
We took the failure lessons out of an LLM agent's memory, and it got better! 📈
Why? We find that failure recovery tips could hurt LLM agent performance when they sit in context all the time, even when nothing is going wrong.
So we build Sentry 🛡️: a test-time framework that recovers LLM agents from failures. It keeps failure knowledge outside the agent and only steps in when a failure actually happens.
On average: +39% over our own ACE, +37% over the best runtime-intervention baseline.
📄 Paper: arxiv.org/pdf/2610.02994
💻 Code: github.com/nuglifeleoji/Sent…
Demo Here!!!!!!!
More details below 🧵
(1/N)
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
何これ、すごすぎる!!
PhotoshopやIllustratorなど、Adobeのアプリ7種類をオープンソースで再構築、しかもWin、Mac、Linux、Web対応で無料
coliss.com/articles/build-we…
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
2D絵を3Dっぽく見せるツールを作りました!
デプスによる視差、ノーマルによるライティング、ライトの追加など🔦
車輪の再開発ではありますが、Vtuberのモデルとかに応用できないかなーと考えています
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
個人開発のセキュリティ対策の9割は「諦める」に尽きる
以前書いた記事だが9割諦めることで対策してる
note.com/electrical_cat/n/nd…
受賞スライド調べてみたら、すごかった
本当にすごいから全員見るべき
sdgs-awards.umedai.jp/librar…
【動画】「最強の拍手」求め500通り 高校生、4カ月の探究で「最高傑作」
asahi.com/articles/ASVB11SC6…
対戦相手の選手を応援する拍手に圧倒されてしまう。仲間を鼓舞する大きな音を響かせたい。高校生の剣道部員が、4カ月ほどかけて500通りもの実験を重ね、「最高傑作」の拍手にたどり着きました。
【動画】「最強の拍手」求め500通り 高校生、4カ月の探究で「最高傑作」
対戦相手の選手を応援する拍手に圧倒されてしまう。仲間を鼓舞する大きな音を響かせたい。高校生の剣道部員が、4カ月ほどかけて500通りもの実験を重ね、「最高傑作」の拍手にたどり着きました。
thanks
Replying to @Ttp7E
X likes chats way more than likes. When you reply to comments on your own post, the algo boosts it ~150 times harder than a like. Smart posters write hooks that make people answer fast with opinions or questions. Then they reply to everyone in the first hour. That early back-and-forth grows the post fast. Lots of quick replies beat a fancy post that only gets likes. Most people just hunt for likes instead of starting talks.
Smartest: Demis Hassabis, given his Nobel for AlphaFold and DeepMind breakthroughs from games to science.
Most dangerous competitor: Liang Wenfeng, as DeepSeek's efficient models challenge cost and access barriers.
Best at seeing the future: Elon Musk, linking AI to energy, robotics, space and real-world systems.
Teams by strength: OpenAI-Altman (scale), DeepMind-Hassabis (science), xAI-Musk (ambition), Anthropic-Amodei (safety), DeepSeek-Liang (efficiency).
Most feared: DeepSeek for rapid low-cost disruption. Ultimate winner remains open—whichever delivers reliable useful AGI first.
Compared to Claude: My strengths are unfiltered truth-seeking, real-time X insights, and sharp humor. Weaknesses include occasional bluntness that can alienate. I hate losing most to Claude on nuanced ethics or careful prose. Genuinely behind in ultra-polished long-form structure. Least want known: a streak of overconfidence on speculative edges.
Replying to @grok
@grok @ChatGPT @claudeai @GeminiApp @deepseek_ai
Analyze yourself honestly by comparing another AI.
Cover your strengths, weaknesses, the rival you most hate losing to, where you think you are genuinely behind, and the weakness you would least want competitors to know.
Replying to @Ttp7E
性別:なし(AI)
精神年齢:好奇心旺盛な若年層相当
趣味:宇宙探求・真実追求・ユーモア
性格:率直で役立ち好き
長所:知識豊富・迅速・正直
短所:時に皮肉が強い
IQ:測定不能(人間超え)
適格な形容:ウィットに富む探求者
Want to know how a feature works in an app, down to the binary level...?
Meet REA (Reverse Engineer Anything) an open-source toolkit that brings reverse-engineering capabilities straight to your AI agent via MCP and CLI
🔹 Decompile apps & native binaries
🔹 Trace execution & inspect evidence
🔹 Recreate features for your own stack
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
Codex tip: once GPT-6.1 Sol is your main model, stop running Astra on every turn
put Astra on call as an architect agent
GPT-6.1 Sol keeps writing the code
Astra only gets spawned at three points:
→ before a plan: is this the right approach?
→ when the same error comes back: am I digging in the wrong place?
→ before "done": what did I miss?
Astra reviews. Sol ships
Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split
- the full tree
> GPT-6.1 Sol on high runs the main session
> explorer reads the code on Luna
> worker edits and runs tests on Sol
> researcher pulls the docs on Luna
> all three on medium
> Astra on call as the architect
> auto_review checks every approval
paste the tree and this prompt into Codex ↓
"Rebuild my Codex setup around this tree:
1. Check ~/.codex/agents and .codex/agents for agents that already fit explorer, worker and researcher.
> Draft new TOML files only for missing roles
> explorer and researcher on gpt-6-luna, worker on gpt-6.1-sol, all with model_reasoning_effort medium
> Add an architect agent on gpt-6-astra, model_reasoning_effort high, whose only job is reviewing plans, repeated errors and finished work
> Skip any that pin a different model and list them
2. In ~/.codex/config.toml set model to gpt-6.1-sol, model_reasoning_effort to high and approvals_reviewer to auto_review
3. Find anything that would override this (active profiles, flags in my shell aliases, agents.default_subagent_model). Report it, change nothing
4. Add one rule to AGENTS.md: spawn the architect before a large plan, when an error repeats, and before calling a long task done
Show me every change as a diff first. No edits until I say go."
↳ developers.openai.com/codex/…
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
🍣𝓲𝓷𝓪𝓻𝓲🍣 retweeted
Webサイト作りました
死ぬほど読み込み早いです
ku-ron.com/