@ActuallyIsaaki
iAccount based inGermany
About this account
- Account based in
- Germany
- Connected via
- Germany App Store
Account-level information from X, not a live location or the device used for a specific post.
ML Researcher | Core contributor to MLX | Violin enthusiast 🎻 | Always coding, sometimes watching Anime.
Germany
Joined February 2022
- Tweets2.1K
- Following858
- Followers2.4K
- Likes67.2K
Pinned Tweet
My name got featured in one of Apples WWDC 2025 session videos 👀🍎
Crazy to think my work made it into a global Apple dev session. I’m proud!!
Thanks @awnihannun @angeloskath @Apple
Gökdeniz Gülmez retweeted
Introducing FLUX 3 Action.
An open weights 7B World Action Model that achieves first place on the RoboLab benchmark.
It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.
FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA.
Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson.
Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next.
FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together.
We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
This and @Nativ_AI live transcription in meetings.
When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
Gökdeniz Gülmez retweeted
Coming to MLX-Audio 🚀
@lllucas let’s go!
When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
Guess whats coming to MLX-LM-LoRA!!!!
github.com/Goekdeniz-Guelmez…
Superintelligence should learn from experience through RL.
Introducing FlashREINFORCE:
Critic-Free, Single-Rollout, Asynchronous RL for Agentic Language Models
Reinforcement Learning Should Do REINFORCE!
github.com/yifanzhang-pro/Fl…
Gökdeniz Gülmez retweeted
Really cool !
Coming to MLX Audio
If not there already
🚀 AuK is officially here. Nano banana🍌 for audio
An open-source foundation model for unified speech generation and editing.
Natural-language instructions + reference audio. One interface.
Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation.
Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions.
Code, weights, and demo are live. Try it and share your feedback.
🤗 Paper & upvote: huggingface.co/papers/2609.0…
⭐ GitHub & star: github.com/Tencent-Hunyuan/A…
Gökdeniz Gülmez retweeted
We’ve had an AI breakthrough at Figure and will be showcasing this tomorrow
Gökdeniz Gülmez retweeted
BREAKING: Graphics benchmarks for Apple's M5 Ultra Mac Studio have LEAKED online!
As it turns out, it's actually FASTER than the Nvidia RTX 5090 Desktop GPU in terms of OpenCL vs Metal
🤯🤯🤯
5090: 350,160
M5 Ultra: 360,019
However, with Vulkan, the 5090 is Faster with 375,290
Gökdeniz Gülmez retweeted
> Apple already offers a tool for efficiently running AI models on Apple hardware called MLX
Let's go!
Apple is plotting a return to selling enterprise servers for the first time since 2011, and is talking to Nvidia about providing the networking tech
Relations between Apple and Nvidia appear to be warming up
theinformation.com/articles/…
Stop typing in meetings. ❌
You look busy and you retain nothing.
@Prince_Canuma does it differently. Watch.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Gökdeniz Gülmez retweeted
Today I had lunch in @Alibaba_Qwen HQ in Hangzhou! With @QwenDevs team we have talked about Qwen, Wan, Qwen Image and ModelScope! So much energy and passion 🚀
Be ready because there is a lot to come! 🤐
It was a pleasure meeting @HaaaaaaydenH in person after interacting on X since January!
JOSIE is a peak gen-z model, exhibit A, keep in mind the reasoning style emerged completely on its own.
Here is the full 105 sample output dataset
huggingface.co/datasets/Goek…
you can inspect the outputs yourself huggingface.co/datasets/Goek…
Nativ v0.3.8 is out🔥🚀
More windows, more room for models, and less friction in every chat.
→ Multiple windows: separate workspaces, each with its own chats, drafts and navigation. File > New Window, or Shift-Command-N
→ External model storage: keep models on an external APFS drive, with connection status and free space shown
→ Quicker chat actions: Return confirms a pending tool action, Up recalls your last prompt in an empty composer, and you can quote selected text in a reply
→ Voice dictation: set your own Return command
→ Smoother long chats: a new Markdown renderer and paged history, plus pinned chats and bulk session controls
→ Better download information: unsupported models flagged, verified sizes, and a disk space check before transferring
→ More connected tools: LWC joins the MCP catalog for read-only project memory, alongside guided DeepSeek Harness setup
→ Settings: opt into beta updates under Advanced, notification controls now live in Permissions
→ More resilient audio capture: recovers when input devices or formats change, and keeps partial recordings
Another step toward making Nativ feel fast, native, and reliable.
Thanks to everyone who contributed🙏
Try Nativ👇:
blaizzy.github.io/nativ/