@bertgodeli
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Pinned Tweet
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions.
We fix that today with OSS Paper Office:
great read.. high-mix, low-volume workloads have created many small precision machine shops
in many ways that diversity is good but, unless you're spacex, it's very cumbersome to scale as an oem
if a product does need to scale, i think redesign for manufacturing probably makes more sense than vertical integration
tho there should also be a deeper network of low-mix, high-volume shops to "graduate to" that are purpose built for longer runs, fewer setups/changeovers, utilization, etc
Thinking about this and concluding that post-training research is now fundamentally a pedagogy problem.
The "post-training data-mixture search space" is too high-dimensional, and therefore techniques and results don't replicate beyond contrived Qwen-8b ablations.
Only thing that works: have an intuition for the model the way a human teacher has one for a human student.
Daanish Khazi retweeted
found out with the @paperinstr team that hash based session routing for agent rl rollouts can deathspiral your samplers and kill your job's throughput.
for max sampling throughput, you generally want to operate close to the kv cache capacity. leaving too much headroom wastes capacity, but crossing the limit causes evicted rollout prefixes to be re-prefilled.
your job's max concurrent rollouts per replica is lower bounded by kv cache capacity:
max rollouts / replica >= kv cache capacity / kv bytes at max context length
in practice, each rollout's working set is usually much smaller than the max context. if average rollout residency is ~50% of max context, you can schedule maybe ~2x that (with headroom)
agent sessions with cumulative context are typically pinned to sampling replicas to maximize prefix cache hits. standard inference engines have no session awareness, so during environment interactions, a rollout's cached kvs become eligible for eviction. if the engine has a request queue from other concurrent rollouts and is under cache pressure, the engine will admit the new prefill and evict the cache of any "idle" rollouts.
hash based routing (e.g. vllm router's consistent_hash policy) is a popular load balancing choice for sampling replicas. hash based routing produces uniform session distribution in expectation, but does not guarantee perfectly balanced live session counts.
skew creates overloaded and underloaded replicas, and an overloaded replica can become "poisoned" with no escape valve. when rollouts on underloaded replicas finish, their replacement can hash to an already overloaded one.
cache pressure -> eviction -> re-prefill -> lower throughput -> bigger queue -> more cache pressure
a simple mitigation is using session-based least loaded assignment. new sessions are assigned to the replica with the least number of active sessions. if you conservatively cap in-flight rollouts to the per replica lower bound described above, least loaded placement entirely prevent eviction. if you intentionally oversaturate based on expected residency, even though eviction is possible, least loaded placement give the useful property that the replica under pressure becomes quarantined. new rollouts stop being scheduled on the replica under pressure until one of its existing sessions finishes, allowing the system to recover.
perfect least loaded session routing does require the router to know when a trajectory is complete so it can decrement its live session count. many generic inference routers don't expose a "release session" api because they are designed around requests, not sessions.
a throughputmaxxing rollout scheduler should dynamically change rollout concurrency, pause rollouts, or explicitly offload/evict cache by session based on kv cache pressure and environment interaction time. this is especially important for rl where trajectories can grow over time as the policy changes, or when tasks have substantially different length distributions.
Testing Stealth Model against GPT-6 Astra, Muse Spark 1.3, and Grok 4.7 one shotting EV drivetrain market study
Same harness, prompts, etc. Crazy results this is white-collar AGI.
Daanish Khazi retweeted
opsd and its variants don’t work, no matter what they are biased and bias kills llms, please please please stop working on it, so many other great directions to explore that aren’t opsd
Meta's team proposes DCE + SRCL as an alternative to on-policy self-distillation (OPSD).
Achieved significant gains: 30.76% avg acc -> 65.97% on Qwen3-8B
arxiv.org/abs/2609.30652
Daanish Khazi retweeted
average OPSD replication effort: "it doesn't really work but wouldn't it be sick if it did?"
Daanish Khazi retweeted
If you’ve built agents or envs that work with word/excel you know the pain of aligning agents to work with docs
Paper is promising as a drop-in replacement to fix a lot of those issues
If you're trying out the packages feel free to use this skills to prime your agent (packages aren't yet in distribution)
github.com/paper-instruments…
Daanish Khazi retweeted
this is absolutely great
Daanish Khazi retweeted
I’ve always complained about how insanely broken AI is for the entire office suite
Happy to see @paperinstr fixing this
Daanish Khazi retweeted
Some very cool shit the @paperinstr team is working on. You should probably check it out.
Daanish Khazi retweeted
this is super helpful for anyone working in legal/professional service domains. shoutout @thegavinbains and team for launching this!
nitter.cf/bertgodel/status/21028…
Daanish Khazi retweeted
it surprised me how often ai agents corrupted @Microsoft office files, even when running simple tasks.
turns out the packages they use haven’t been updated in years and in some cases fully rewrite a file’s XML on save, destroying or invalidating special features like charts.
paper office is part of our efforts to improve agents’ abilities to create/edit and maintain these files
Daanish Khazi retweeted
Your agents will be much happier when you use @paperinstr 📄