@byte_lab

Ring 0 Linux kernel hacker at Meta, scheduler-bug-adder, BPF standardization co-chair. If you want to know how it works, break it apart.

Chicago, IL
Joined March 2020
David Vernet retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,654
16,389
11,352
87,840
76,492,191
I guess now is as good a time as any to announce that I’m starting at Anthropic tomorrow. Super grateful for the opportunity!
9
79
5,168
It’s interesting how one’s relationship to music evolves as you get older. Thirty Three has always been one of my favorite @SmashingPumpkin song, but now I love it primarily because of the lyrics. “Speak to me in a language I can hear.” Now that I’m older, I can really hear it.
3
585
Proud to see this go into the standard, and humbled to be a C++ standards coauthor. Please don’t blame me for the rest of the language, though.
The Linux kernel's most important concurrency primitive has been in production for 20+ years Readers never block. Writers make a copy, modify the copy, then atomically swap the pointer C++26 standardizes it as std::rcu 🧵👇
1
7
1,618
Woah indeed, my friend
Woah! @AMD P-State Linux Driver Patches Can Boost 1%-Low FPS Gaming Performance By 31% Wild improvement for the Steam Deck and other modern @AMDRyzen hardware benefits too with new proposed "epp_boost" Linux feature for amd_pstate driver. phoronix.com/news/AMD-P-Stat…
2
13
1,664
@gautshen accidentally cc’d your old AMD email address. Sorry, old habits die hard.
1
101
Random friends and family have asked me over the last few months and years what I think about immigration in the US. I'll answer by saying this: there is a permanent exhibit on display at Ellis Island that plays a recording of an interview of my great grandmother who immigrated here from Poland to flee antisemitism. She was recalling one of the logic questions she was asked when making her way through immigration at the island: Their question: Do you wash stairs starting at the top, or starting at the bottom? Her answer: I didn't come to America to wash stairs That is the core of the American spirit. You have the best people coming to the best place. Anecdotally, every single Indian and Chinese person I've met in the US (literally every single one) are highly skilled laborers who pay an exorbitant tax bill directly to the US. The second you turn off that tap, this country just becomes another gaggle of homogeneous underachievers, and I for one won't stand for it. There is neither a cogent economic, nor moral reason to be against legal immigration. You get either skilled labor that pours fuel on American industry (thanks for making me rich, tech industry), or unskilled labor that drives down costs and raises standards of living.
3
380
Another area I was snooping recently is in the AMD cpufreq driver. There's a common pattern in gaming where you have a main task and a render task that are (largely) serialized on the rendering pipeline; typically with the rendering task blocking on some futex while the GPU does its thing. You usually always want to jack the frequency up to performance levels on whatever cores they're running on if you can, because these tasks are essentially the embodiment of Amdahl's Law when it comes to frame rendering. In other words, they're serialized and on the critical path for rendering, so if either of them slow down, the entire rendering pipeline is now more likely to result in a stale frame (i.e. a frame being delivered to the compositor after the frame deadline). The AMD cpufreq governor kind of works against you here, because these tasks are short and bursty, and the firmware cpufreq governor seems to decay its tracked util between frames to the point that, at the start of the next frame when the core is no longer idle, it starts at a low frequency and has to ramp back up to the higher frequency over the course of the frame. The solution, I think, is to have some per-core utilization tracking that writes EPP=performance to MSR_AMD_CPPC_REQ (by default on e.g. the Steam Deck it's EPP=balance_performance) on some sampling cadence if util is high enough over that period, and then drops it back to whatever it was before when the core goes idle. I tested a local patchset and in the Civ 6 graphics benchmark (on the Steam Deck) it improved 1%-low fps by 31.8% (which is fairly massive), and p99 frame time by 4.1%. Hoping to polish them and send them upstream in the next day or two.
2
1
6
1,760
Just noticed yesterday that the first 4 fds (on 64 bit) in the `fd_array` in `struct files_struct` (git.kernel.org/pub/scm/linux…) share the same cacheline as the write-heavy fields like `file_lock`. This means that if multiple threads are doing write operations on fds, it can cause false sharing for fds 0-3. Is this intentional? Looks like it's been this way since ~2005 so I assume so, but I also wonder if we should revisit this now that multi-processor systems are the norm. I tested it and if you cache-line align `fd_array`, the whole struct still fits into 11 cache lines.
2
10
1,858
Ah, this was previously discussed in commit 0c9e63fd38a2 ("[PATCH] Shrinks sizeof(files_struct) and better layout") bit.ly/4yvQHfK: >UP and SMP should benefit from this patch, because most tasks will touch only one cache line when open()/close() stdin/stdout/stderr (0/1/2), (next_fd, close_on_exec_init, open_fds_init, fd_array[0 .. 2] being in the same cache line) I kind of doubt that holds true today in general.
2
150
As per usual, the Intel variant is far more opaque and stateful. You can only touch the VMCS with vmread and vmwrite instructions. AMD uses the VMCB which is just in RAM and can be accessed with normal loads and stores.
interested in architectural differences between svm and intel's vmx. most of the features seem to be Intuitively similar
4
1,655
In my opinion the core problem with EEVDF is that it uses slice length as a knob for latency. This is a mistake — like existing interfaces in the POSIX scheduling API such as niceness, it is an extremely awkward API to use correctly. How do you determine how long a task should be scheduled for to keep its latency low, but not SO low that context switch overhead becomes a problem? It’s literally impossible to reason about this for even the world’s foremost scheduling experts. The only usable interface for this that I’ve seen is QoS based, and the only framework that realistically enables this is sched_ext.
1
13
740
Still my favorite bug I’ve ever found (with colleagues): an Apple vulnerability where they didn’t issue an LGDT after a vmexit. You have to because the VMCS will switch the GDT base address, but not the size of the GDT, and if you don’t, you can end up with an arbitrarily large GDT. This means you just have to find a descriptor you want and you can start cranking your code at CPL 0. I don’t blame them — what kind of psycho designs an interface that replaces the base address of a critical hardware data structure, but not the size (@intel does)?
1
4
157
Many people in the tech industry (including, apparently, some AWS engineers and Andy Jassy himself) seem to have a fundamental misunderstanding about "security". Security "vulnerabilities" are simply bugs. They are not all necessarily emergencies, and they are not all created equal. Bugs are also sometimes not security vulnerabilities. For example, this code of returning an address on the stack as a pointer in a function call is a classic mistake that many new C programmers make: ``` static int *my_func(void) { int my_var = 5; return &my_var; } ``` This is a bug because it's completely broken and will likely crash the program, but it can also cause all sorts of memory corruption issues that could be taken advantage of by an attacker. If Fable 5 found this "vulnerability", it would not be bypassing some sanctified security guardrail. It's just a bug in software that needs to be fixed. There is no evidence whatsoever that I'm aware of that anyone has found a way to bypass the guardrails that Anthropic baked into the model. Finding a bug != bypassing the security guardrails. If that were the case then every model including Haiku would have to be taken off the market. >if it wasn't for the classifier i can empty your bank account with fable/mythos. You clearly didn't even read their post. There is no way that anyone who knows anything about software could have read the latest Anthropic news post and come away with any conclusion other than that this whole thing was total BS.
Replying to @byte_lab
again. its not about finding vulnerabilities or what model found the same vulnerabilities, its about bypassing fables classifier TO FIND VULNERABILITIES like i don't understand.. do you not get the difference between haiku and fable/mythos class models? if it wasn't for the classifier i can empty your bank account with fable/mythos. you can not do that with opus, or gpt5.5 let alone haiku
1
2
157
Whomever the engineers were at Amazon who submitted this "jailbreak" report -- I hope you have some ability to feel shame. Imagine escalating this report all the way to the *White House* when Haiku 4.5 identified the same """"vulnerabilities"""" that you found with Fable 5. Congratulations on accomplishing literally nothing other than pissing everybody off for 2 weeks.
Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests. We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort. Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research. Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again. Read our full blog: anthropic.com/news/redeployi…
19
8
342
28,214