@bertgodel
sf
Joined February 2018
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
42
53
16
774
74,485
great read.. high-mix, low-volume workloads have created many small precision machine shops in many ways that diversity is good but, unless you're spacex, it's very cumbersome to scale as an oem if a product does need to scale, i think redesign for manufacturing probably makes more sense than vertical integration tho there should also be a deeper network of low-mix, high-volume shops to "graduate to" that are purpose built for longer runs, fewer setups/changeovers, utilization, etc
1
6
498
Thinking about this and concluding that post-training research is now fundamentally a pedagogy problem. The "post-training data-mixture search space" is too high-dimensional, and therefore techniques and results don't replicate beyond contrived Qwen-8b ablations. Only thing that works: have an intuition for the model the way a human teacher has one for a human student.
age of research for post-training ended a few months ago: "the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training"
3
1
3
60
6,384
Daanish Khazi retweeted
found out with the @paperinstr team that hash based session routing for agent rl rollouts can deathspiral your samplers and kill your job's throughput. for max sampling throughput, you generally want to operate close to the kv cache capacity. leaving too much headroom wastes capacity, but crossing the limit causes evicted rollout prefixes to be re-prefilled. your job's max concurrent rollouts per replica is lower bounded by kv cache capacity: max rollouts / replica >= kv cache capacity / kv bytes at max context length in practice, each rollout's working set is usually much smaller than the max context. if average rollout residency is ~50% of max context, you can schedule maybe ~2x that (with headroom) agent sessions with cumulative context are typically pinned to sampling replicas to maximize prefix cache hits. standard inference engines have no session awareness, so during environment interactions, a rollout's cached kvs become eligible for eviction. if the engine has a request queue from other concurrent rollouts and is under cache pressure, the engine will admit the new prefill and evict the cache of any "idle" rollouts. hash based routing (e.g. vllm router's consistent_hash policy) is a popular load balancing choice for sampling replicas. hash based routing produces uniform session distribution in expectation, but does not guarantee perfectly balanced live session counts. skew creates overloaded and underloaded replicas, and an overloaded replica can become "poisoned" with no escape valve. when rollouts on underloaded replicas finish, their replacement can hash to an already overloaded one. cache pressure -> eviction -> re-prefill -> lower throughput -> bigger queue -> more cache pressure a simple mitigation is using session-based least loaded assignment. new sessions are assigned to the replica with the least number of active sessions. if you conservatively cap in-flight rollouts to the per replica lower bound described above, least loaded placement entirely prevent eviction. if you intentionally oversaturate based on expected residency, even though eviction is possible, least loaded placement give the useful property that the replica under pressure becomes quarantined. new rollouts stop being scheduled on the replica under pressure until one of its existing sessions finishes, allowing the system to recover. perfect least loaded session routing does require the router to know when a trajectory is complete so it can decrement its live session count. many generic inference routers don't expose a "release session" api because they are designed around requests, not sessions. a throughputmaxxing rollout scheduler should dynamically change rollout concurrency, pause rollouts, or explicitly offload/evict cache by session based on kv cache pressure and environment interaction time. this is especially important for rl where trajectories can grow over time as the policy changes, or when tasks have substantially different length distributions.
5
9
460
Testing Stealth Model against GPT-6 Astra, Muse Spark 1.3, and Grok 4.7 one shotting EV drivetrain market study Same harness, prompts, etc. Crazy results this is white-collar AGI.
3
23
1,183
Open source needs to start training with apply_patch
Enterprise teams can now use open models like GLM-5.3 Flash and Kimi K3 natively in Codex and count spend against their OpenAI commit.
1
6
603
Daanish Khazi retweeted
opsd and its variants don’t work, no matter what they are biased and bias kills llms, please please please stop working on it, so many other great directions to explore that aren’t opsd
Meta's team proposes DCE + SRCL as an alternative to on-policy self-distillation (OPSD). Achieved significant gains: 30.76% avg acc -> 65.97% on Qwen3-8B arxiv.org/abs/2609.30652
24
33
11
436
86,636
Daanish Khazi retweeted
average OPSD replication effort: "it doesn't really work but wouldn't it be sick if it did?"
12
8
1
240
16,231
this is da agi
4
13
1,097
"do you want to pace the frontier?" "no, do you?" "no"
1
18
647
Daanish Khazi retweeted
If you’ve built agents or envs that work with word/excel you know the pain of aligning agents to work with docs Paper is promising as a drop-in replacement to fix a lot of those issues
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
1
7
1,835
Daanish Khazi retweeted
I’ve always complained about how insanely broken AI is for the entire office suite Happy to see @paperinstr fixing this
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
1
12
835
Daanish Khazi retweeted
Some very cool shit the @paperinstr team is working on. You should probably check it out.
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
1
9
1,128
Daanish Khazi retweeted
Let’s go!
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
1
9
3,181
Daanish Khazi retweeted
this is super helpful for anyone working in legal/professional service domains. shoutout @thegavinbains and team for launching this! nitter.cf/bertgodel/status/21028…
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
2
12
629
Daanish Khazi retweeted
it surprised me how often ai agents corrupted @Microsoft office files, even when running simple tasks. turns out the packages they use haven’t been updated in years and in some cases fully rewrite a file’s XML on save, destroying or invalidating special features like charts. paper office is part of our efforts to improve agents’ abilities to create/edit and maintain these files
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
2
3
15
481
Daanish Khazi retweeted
Your agents will be much happier when you use @paperinstr 📄
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions. We fix that today with OSS Paper Office:
1
8
734