AMUN REIGN retweeted
HuggingFace just closed a major gap in harness engineering!
(the ultimate guide to multi-harness RL)
the same open-weight model can perform well inside one harness, then lose accuracy or produce invalid tool calls when moved to another.
this happens because training inside one interface can teach the model its specific tool names, output formats, context structure, and control flow. the model learns how to operate the harness, not just how to solve the task.
Hugging Face researchers tested a more portable approach called multi-harness reinforcement learning. instead of training a model through one agent interface, they trained it through Claude Code, Codex, OpenCode, and Mini-SWE-Agent.
to make that possible, they connected three open systems.
โ ๐ข๐ฝ๐ฒ๐ป๐๐ป๐ provides a standard interface between agent harnesses, reinforcement learning environments, and trainers. its capture proxy sits between the harness and model server, recording the exact tokens and generation probabilities needed for training.
โ ๐๐ฎ๐ฟ๐ฏ๐ผ๐ฟ runs agents against containerized tasks. it keeps the task, harness, and sandbox independent, allowing the same task to be attempted through different harnesses without rebuilding the environment.
โ ๐ง๐ฅ๐ is Hugging Faceโs open-source library for post-training language models. it uses the trajectories captured by ๐ข๐ฝ๐ฒ๐ป๐๐ป๐ to update the model through reinforcement learning.
the real breakthrough is where the training data gets captured.
each harness keeps its native tools, prompts, context management, retries, and execution loop. the researchers do not recreate those behaviors inside the trainer. ๐ข๐ฝ๐ฒ๐ป๐๐ป๐ observes the model calls passing through each real harness and converts them into usable training sequences.
the team trained ๐๐๐ ๐ฎ.๐ฑ-๐ฎ.๐ฒ๐ across all four harnesses. the share of held-out tasks solved on the first attempt increased from 42.2% to 54.2%, with gains under every harness.
the trained model also used 31% fewer tool calls on tasks that both it and the base model solved.
training only in OpenCode improved the model too, but most of that gain stayed concentrated in OpenCode. multi-harness training spread the improvement across interfaces.
the experiment used one task family, one training seed, and unequal data exposure, so the results are not a universal ranking. even with those limitations, the mechanism matters.
open-source models cannot assume one deployment interface. if they need to work across harnesses, that portability must become part of training.
Read the full guide here: huggingface.co/spaces/FineEnโฆ
if you want to understand harness engineering and what an agent harness actually includes, i wrote a full breakdown to help you get started. the article is quoted below.
AMUN REIGN retweeted
sam altman revealed in 2024 that GPT-6 took 9 years and $87B to train - and 99,8% of people still don't know what actually happened inside that model.
a 2026 working note just dropped that maps the exact loss landscape of GPT-6. and the numbers are insane.
~1T parameters. 200K training steps. min loss: 1.76. generalization gap: 0.28 - the lowest ever recorded across any frontier model.
not the architecture. not the data. the geometry.
for comparison: Llama 3 at 8B sits at 0.42. Claude 3.5 at 200B gets to 0.33. GPT-6 at ~1T hits 0.28. scale doesn't just improve performance - it changes the shape of the entire loss surface.
here's what that means in practice:
the loss landscape L(ฮธ) is a high-dimensional surface with saddle points, sharp minima, and flat basins. where your model lands after training determines everything - generalization, robustness, emergent capabilities.
sharp minima = model memorizes. overfits. collapses in production.
flat basins = model learns structure. generalizes. scales. this is where GPT-6 lands after AdamW optimization across 200K steps.
to measure this precisely they used Hessian analysis. H(ฮธ) = โยฒL(ฮธ). GPT-6's eigenvalue spectrum is heavy-tailed - a few large eigenvalues, thousands near zero. translation: the optimization surface is smooth at scale. smaller models don't get this property. they stay stuck in narrow valleys.
this is why emergent capabilities - in-context learning, chain-of-thought, tool use - aren't engineered. they appear when the geometry becomes navigable enough for the optimizer to escape local minima and settle into flat basins.
$100B. 9 years. 200K steps. the breakthrough wasn't more compute.
it was understanding the shape of the space the model lives in.
Bookmark & watch today.
know the geometry or compete against people who do.
AMUN REIGN retweeted
OpenAI Dots is f*cking brilliant for building a 24/7 AI company...
hereโs the bigger architecture Iโd build around it:
one founder. one dot powered by GPT-6 Astra. 15 specialist jobs connected by work packets
the system has four loops:
BUILD
feedback + support repros โ verified evidence โ scoped spec โ code branch โ tested PR
LAUNCH
approved changes โ explainers + demo clips + documentation โ launch pack
REVENUE
account context โ working POC โ proposal โ follow-up draft
objections and missing features go back into research
OPERATIONS
support triage + invoice drafts + status tracking โ one queue of decisions for the founder
the connections are where this gets useful:
a support ticket can become a repro, a patch, updated docs and an answer draft
a finished feature becomes both launch material and proof for the next proposal
a sales objection becomes evidence for the next product decision
give every handoff a file:
โ source references
โ the actual output
โ checks run + open blockers
โ the next job and its exact context
the dot routes the work. specialists return artifacts. you review the decisions and feed corrections into the next task
start with one loop. make it work. connect the next one
save this, then build your company with Dots โญฃ
AMUN REIGN retweeted
Agents without memory aren't agents at all.
An LLM can appear to remember because the application keeps sending previous messages back with each request.
The model itself is still stateless.
Start a new session without stored history, and every preference, decision, and previous outcome disappears.
Agent memory solves this at two different scopes:
1๏ธโฃ Short-term memory
This is the agent's working state during the current session. It includes recent messages, retrieved documents, uploaded files, tool outputs, and intermediate task state.
2๏ธโฃ Long-term memory
This persists across sessions. It stores information the agent may need again, such as user preferences, known facts, previous outcomes, and task instructions.
Long-term memory can be divided further:
1) Semantic memory stores facts.
For example, a support agent might remember which plan a customer uses or that the customer prefers email over phone calls.
2) Episodic memory stores previous experiences.
This could include how an earlier support issue was resolved or what happened during a previous agent run.
3) Procedural memory stores instructions and learned procedures.
This includes behavioral rules, tool preferences, and steps the agent should follow when completing a task.
These memory types need different write and retrieval policies. A preference should persist until it changes, while an old tool output may only matter during the current session.
The model is not learning through weight updates here. The surrounding system adapts by storing, updating, and retrieving state.
Oracle AI Agent Memory implements this through short-term threads, summaries, durable memories, automatic extraction, scoped retrieval, and context cards.
In Oracle's documented 80-turn evaluation, the system held input near 1,300 tokens per request while flat history grew past 13,900 tokens by the final turn.
The managed-memory agent also won 48 evaluated turns, while flat history won 13. The remaining 19 were ties.
I worked with the Oracle team on this post to explain how the memory architecture works. The visual above breaks down the memory types the system needs to manage.
You can read the full architecture and evaluation here: fandf.co/4AFP7sz
Memory is one slot in a much larger harness. The orchestration loop, the tools, the state persistence, and the guardrails all sit around the model too, and memory is the part that decides what the rest of them see on the next turn.
I broke down all eleven components of a production agent harness in an article.
Read it below.
AMUN REIGN retweeted
Azure is becoming the operating system for enterprise AI.
It is no longer only about hosting models or adding a chatbot to an existing application. Azure AI Foundry brings the full AI development lifecycle into one connected ecosystem.
From choosing the right model to deploying agents, monitoring performance, and securing production workloads, every layer can work together.
The ecosystem is built around four key areas:
1. Design with the best models
Teams can access Azure OpenAI, Phi, DeepSeek, Meta Llama, Mistral, Cohere, Hugging Face, Nvidia, Databricks, Snowflake, and other model providers.
2. Customize with an agent toolchain
Developers can connect agents with Azure AI Search, Fabric, SQL, Cosmos DB, Functions, Kubernetes, Semantic Kernel, LangChain, LlamaIndex, AutoGen, and many other services.
3. Manage production performance
Azure Monitor, App Configuration, Microsoft Cost Management, ClearML, Dataloop, and GitHub Actions help teams observe, optimize, and improve AI systems.
4. Safeguard with trustworthy AI
Content Safety, Microsoft Defender, Entra ID, Azure Policy, Confidential Computing, Backup, and Application Gateway support more secure and governed deployments.
Copilot Studio, Visual Studio, GitHub, and the Azure AI Foundry SDK connect the entire development experience.
The real advantage is not access to more tools.
It is having one ecosystem that supports AI from the first prototype to secure production.
Save this if you are building AI agents on Azure.
๐๐ฒ๐ฐ๐ผ๐บ๐ฒ ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐ฎ๐ ๐๐ ๐ถ๐ป ๐ท๐๐๐ ๐ญ ๐บ๐ถ๐ป๐๐๐ฒ ๐ฎ ๐ฑ๐ฎ๐. ๐๐ผ๐ถ๐ป ๐บ๐ ๐๐ฒ๐ฒ๐ธ๐น๐ ๐ป๐ฒ๐๐๐น๐ฒ๐๐๐ฒ๐ฟ ๐๐ต๐ฒ๐ฟ๐ฒ ๐ ๐ฑ๐ผ๐ฐ๐๐บ๐ฒ๐ป๐ ๐๐ต๐ฒ ๐ฟ๐ฒ๐ฎ๐น-๐๐ผ๐ฟ๐น๐ฑ ๐ท๐ผ๐๐ฟ๐ป๐ฒ๐ ๐ผ๐ณ ๐๐ ๐๐ฟ๐ฎ๐ป๐๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป.
๐ ๐ฆ๐ถ๐ด๐ป ๐๐ฝ ๐ณ๐ฟ๐ฒ๐ฒ now โ avsl.beehiiv.com/
Follow @AiswaryaVenkit1 for more such insights!!
AMUN REIGN retweeted
Jev + LLMs is f*cking insane...
this free Jev lab is a clean blueprint for agent builders
one support ticket enters. Jev reads the state and answers three typed questions in one call:
noul โ is it urgent?
choice โ which team owns it?
score โ how soon does it need action?
then code reads the probabilities. clear, low-risk work flows to the right LLM. ambiguous tickets go to a person. refunds wait for approval. the draft gets checked against policy before it goes out
the loop:
ticket โ typed questions โ probabilities โ threshold โ LLM draft โ approval / reply check โ shadow run
the best part of the lab: you can break the system yourself. set the threshold to zero, remove the โotherโ choice, watch routing fail, then repair it. the final exercise runs Jev beside existing rules on 40 tickets without letting it act
the practical rule: Jev for bounded judgments, an LLM for writing, plain code for hard constraints
save this, then read the article below to combine Jev with Dots โญฃ
AMUN REIGN retweeted
๐จAI Agents are getting identities now.
And that changes everything.
Microsoft just introduced Entra Agent ID โ bringing identity, governance, and security controls to AI agents across the enterprise ecosystem.
From authentication to visibility, organizations can now manage AI agents with the same level of trust and control as human users.
Hereโs what stands out ๐
๐น Authentication for AI agents
๐น Authorization & access governance
๐น Identity protection mechanisms
๐น Better visibility into agent activities
๐น Integration across Copilot Studio, Foundry, Microsoft 365 Copilot & third-party ecosystems
This is a major step toward secure enterprise AI adoption.
As AI agents become part of daily workflows, identity management will become non-negotiable.
The future isnโt just AI-powered.
Itโs AI-governed. โก
What are your thoughts on AI identity management?
๐๐ฒ๐ฐ๐ผ๐บ๐ฒ ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐ฎ๐ ๐๐ ๐ถ๐ป ๐ท๐๐๐ ๐ญ ๐บ๐ถ๐ป๐๐๐ฒ ๐ฎ ๐ฑ๐ฎ๐. ๐๐ผ๐ถ๐ป ๐บ๐ ๐๐ฒ๐ฒ๐ธ๐น๐ ๐ป๐ฒ๐๐๐น๐ฒ๐๐๐ฒ๐ฟ ๐๐ต๐ฒ๐ฟ๐ฒ ๐ ๐ฑ๐ผ๐ฐ๐๐บ๐ฒ๐ป๐ ๐๐ต๐ฒ ๐ฟ๐ฒ๐ฎ๐น-๐๐ผ๐ฟ๐น๐ฑ ๐ท๐ผ๐๐ฟ๐ป๐ฒ๐ ๐ผ๐ณ ๐๐ ๐๐ฟ๐ฎ๐ป๐๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป.
๐ ๐ฆ๐ถ๐ด๐ป ๐๐ฝ ๐ณ๐ฟ๐ฒ๐ฒ now โ avsl.beehiiv.com/
Follow @AiswaryaVenkit1 for more such insights!!
AMUN REIGN retweeted
Everyone wants to build their own company brain.
Gorgias actually did, with a team of four in about 3 to 4 weeks.
(the year before that is the part that mattered most for them)
Yochan Khoi opened up Cortex, the company brain Gorgias runs on, right next to Slite's. @Christophepas asked him the question every team is weighing: should you build your own?
What I took away:
/1 The build is the quick part
Gorgias spent about a year on another tool before they knew exactly what they wanted. Today, roughly half of Yochan's time still goes to new features and changes. Most of the rest goes to helping people use it well.
That upkeep is the part most teams underestimate. It's also what we've been building maintenance for at Slite.
/2 Seed it like you're onboarding someone
Every department wrote down what they'd tell a new joiner who wasn't allowed to talk to anyone for a week. Yochan's reasoning: clever chunking only matters when you don't control how knowledge gets written, and when you do, the structure itself is what makes answers accurate.
/3 Ownership is the whole point
Anyone can request a change, but an owner approves it before it goes live. If an answer is wrong, you know who to ask. Think of it as pull requests for your knowledge.
/4 Business teams use it most
Most questions come from support, CX and sales. Engineers use it the least, since their answers usually already live in the code.
So, should you build?
Yochan's advice: if you're a small company, or you've never shipped agentic systems or internal apps, don't. If you have the people and the experience, it can be worth it. For Gorgias it was a win-win: they sell AI agents, so everything they learn feeds the product.
Which part would you still build yourself?
Watch the full session and read the recap here: slite.com/webinars/company-bโฆ
AMUN REIGN retweeted
RAG vs. CAG, clearly explained!
In a standard RAG setup, every query hits the vector DB, including queries about a product manual or policy documents that haven't changed in months.
The retrieval adds latency, and then the model prefills those same retrieved chunks again on every subsequent query.
CAG is a technique that drops the vector search and moves the prefill off the query path.
The preprocessing step runs those documents through the model once, before any query arrives, and keeps the key and value tensors it produces for every token at every layer.
At query time, the model loads that state and starts decoding, with no vector search or prefill on the knowledge.
The amount of context you can store as cache isn't limited by the context length of the model but rather the GPU memory.
For instance, in a 70B model at BF16, the cache takes around 300 KB/token, so even a small corpus can produce tens of GBs of cache to manage.
That's why production setups run both RAG and CAG together.
โณ Static, high-value knowledge that nearly every query reads gets cached once, like policies, product docs, and standing instructions.
โณ Everything else stays in the vector DB, since a document that surfaces in one query out of a thousand doesn't justify holding its tensors on the GPU all day.
The diagram below depicts this.
To use this in practice, you don't need to build a custom serving stack.
The transformers library already implements the cache as an object of KV vectors that you can preserve, so you can prefill a corpus once, retain the returned tensors, and reuse them across queries in about ten lines.
And this KV cache is only one of four separate caching layers in an LLM stack.
The other three are prefix caching on the server, prompt caching billed by a provider, and a semantic cache that skips the model entirely.
We wrote a full breakdown of all four caches in LLM serving that you should know as an AI engineer, with code for each.
Read it below.
AMUN REIGN retweeted
Jev for RAG, clearly explained!
Hybrid search gives you a shortlist. It does not decide which passages contain evidence, which are merely adjacent, or whether the evidence is strong enough to answer.
That missing judgment is where Jev fits.
Jev sits between retrieval and generation. It does not replace BM25, embeddings, or the LLM. It evaluates the candidates they produce before those candidates enter the context window.
โ Retrieve wide
Combine dense and keyword search, then merge the results with reciprocal rank fusion. This gives Jev a broad candidate pool, such as the top 20 passages shown in the visual.
Retrieval still sets the ceiling. If the right passage is missing from this shortlist, Jev cannot recover it.
โ Judge every candidate together
Send the query as Jevโs state. For each candidate, ask a typed yes-or-no question such as โDoes C7 help answer this query?โ
Jev evaluates every question in one packed request and returns a calibrated probability for each passage. This avoids making a separate model call for every query-passage pair.
โ Let code apply the threshold
Your application compares each probability with a threshold. Candidates above it continue to the LLM. Everything below it is removed before generation.
Jev makes the fuzzy judgment. Code remains responsible for the actual decision.
โ Gate the entire answer
The same request can check whether the retained passages make the query answerable. It can also flag signs of prompt injection inside a candidate.
If answerability falls below the threshold, the application can skip the LLM and return โnot in the documents.โ The injection score should remain a filtering signal, not a security boundary.
The result is a cleaner division of work.
Hybrid search retrieves broadly. Jev reranks, filters, and decides whether sufficient evidence exists. The LLM writes only from the passages that survive.
Jevโs value here is not simply moving passages up or down a list. It turns relevance into an explicit probability that your application can inspect, threshold, and act on.
To summarise:
- Retrieval finds the candidates.
- Jev decides what deserves context.
- The LLM writes the grounded answer.
----
I also built an open-source project showing how to use Jev as a judge for AI observability with Comet Opik.
It evaluates support traces for groundedness, request coverage, action honesty, and helpfulness, then records the results as an auditable experiment.
You can explore the project here: github.com/patchy631/jev-as-โฆ
My article on how Jev works is quoted below.
AMUN REIGN retweeted
๐ง๐๐ ๐๐จ๐๐ ๐ฆ๐ง๐๐๐ ๐๐๐๐๐ก๐ ๐๐๐๐ก๐ง๐๐ ๐๐
Building an AI agent is not just about choosing an LLM.
A production-ready agentic AI system needs multiple layers working together:
01 โ ๐๐ฅ๐ข๐ก๐ง๐๐ก๐
The user-facing layer for interacting with the AI.
Tools:
React, Next.js, Streamlit, Azure App Service
02 โ ๐๐ข๐๐จ๐ ๐๐ก๐ง ๐๐ก๐๐๐ฆ๐ง๐๐ข๐ก
Bring data from documents and other sources into the system.
Tools:
Azure AI Content Understanding, Apache Tika, Microsoft Fabric, LangChain
03 โ ๐๐๐จ๐ก๐๐๐ก๐ & ๐ฃ๐ฅ๐๐ฃ๐ฅ๐ข๐๐๐ฆ๐ฆ๐๐ก๐
Break large documents into useful, searchable pieces before sending them to the model.
Tools:
spaCy, Hugging Face, LangChain
04 โ ๐๐ ๐๐๐๐๐๐ก๐๐ฆ
Convert text into vectors so the system can understand semantic relationships.
Tools:
OpenAI, Cohere, Azure AI
05 โ ๐ฉ๐๐๐ง๐ข๐ฅ ๐๐๐ง๐๐๐๐ฆ๐
Store and search those embeddings efficiently.
Tools:
Azure Cosmos DB, Azure PostgreSQL, Milvus, FAISS
06 โ ๐ฅ๐๐ง๐ฅ๐๐๐ฉ๐๐ ๐๐๐ฌ๐๐ฅ
Find the most relevant information before generating an answer.
Tools:
Azure AI Search, LangChain, LlamaIndex
07 โ ๐ฃ๐ฅ๐ข๐ ๐ฃ๐ง ๐๐ก๐๐๐ก๐๐๐ฅ๐๐ก๐
Turn retrieved context into effective instructions for the model.
Tools:
Promptify, LangChain, DSPy
08 โ ๐๐๐
The intelligence layer that reasons over the provided context.
Examples:
Azure AI, OpenAI, Llama, Mistral AI
09 โ ๐๐ก๐๐ฅ๐ / ๐๐๐ฃ๐๐ข๐ฌ๐ ๐๐ก๐ง
Run and scale the AI application reliably.
Tools:
Azure Container Apps, AKS, Docker, Kubernetes
10 โ ๐ข๐๐ฆ๐๐ฅ๐ฉ๐๐๐๐๐๐ง๐ฌ & ๐๐ฉ๐๐๐จ๐๐ง๐๐ข๐ก
Monitor performance, trace workflows and evaluate outputs.
Tools:
Azure Foundry, OpenTelemetry, Grafana
๐ง๐๐ ๐๐๐ ๐๐๐๐:
Agentic AI is not one model.
It's a complete pipeline:
๐๐ฎ๐๐ฎ โ ๐๐ต๐๐ป๐ธ๐ถ๐ป๐ด โ ๐๐บ๐ฏ๐ฒ๐ฑ๐ฑ๐ถ๐ป๐ด๐ โ ๐ฅ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น โ ๐ฃ๐ฟ๐ผ๐บ๐ฝ๐๐ โ ๐๐๐ โ ๐๐ฒ๐ฝ๐น๐ผ๐๐บ๐ฒ๐ป๐ โ ๐๐๐ฎ๐น๐๐ฎ๐๐ถ๐ผ๐ป
Save this as a roadmap if you're learning AI engineering or building RAG/agentic AI systems.
Repost if this helped you understand the AI stack.
Follow @AamirAnsar94694 for more AI, tools, productivity & tech insights.
#AI #AgenticAI #AIAgents #AIEngineering #RAG #LLM #GenerativeAI #MachineLearning
AMUN REIGN retweeted
Skill Contract Test Suite
A skill with no contract will drift the first time two teams share it.
As a dev, I now test each skill against a written contract.
Contract Tests Cheatsheet:
1. Name required inputs
2. Name allowed outputs
3. Name required refusals
4. Fail CI when the contract breaks
5. Version the contract
6. Review consumers before a breaking change
Core principle: The contract is the API. The prompt is the implementation.
Pro tip: Refusal cases belong in the suite. Happy path only is a toy test.
Does your skill have a test that proves it will say no?
Reply below
Follow @AiCamila_ for daily production AI and DevOps patterns.
#Contracts #Testing #AgenticAI #ProductionAI #Skills
AMUN REIGN retweeted
Skill Scaffold Template
A blank folder is how every skill invents its own layout and then breaks CI.
As a dev, I now start every skill from one scaffold.
Scaffold Cheatsheet:
1. Generate the folder from a template
2. Include a contract file
3. Include fixture tests
4. Include a README stub
5. Include a default cheap model pin
6. Ban skills that skip the scaffold
Core principle: Same shape. Faster review. Fewer missing files.
Pro tip: The scaffold should fail CI if the owner field is empty.
How long does a new skill in your repo take before it has tests?
Reply below
Follow @AiCamila_ for daily production AI and DevOps patterns.
#Scaffold #DX #AgenticAI #ProductionAI #Skills
AMUN REIGN retweeted
holy sh*t, someone mapped an entire dev team into 160+ Claude Code subagents
every AI builder can steal the setup
this repo covers 10 categories across coding, infrastructure, QA, security, data, product and research
the patterns:
โ let api-designer define the contract
โ have backend-developer build it
โ send the changes to code-reviewer
โ run qa-expert and security-auditor before release
โ bring in devops-engineer to prepare deployment
โ use task-distributor for the next handoff
the architecture:
isolated context โ scoped tools โ role-specific model โ clear handoff
reviewers get read-only access. builders get editing tools. smaller tasks can use lighter models.
the part builders should steal:
scope โ build โ review โ test โ secure โ ship
one specialist with a clear job at every step
save this, then build a second brain for your businessโญฃ
AMUN REIGN retweeted
๐ ๐๐ฃ ๐๐ ๐๐ฃ๐.
An ๐๐ฃ๐ defines how software systems communicate through specific endpoints, requests, and responses. It gives applications a structured way to access data or trigger functionality in another system.
๐ ๐๐ฃ gives AI applications a standardized way to discover and use external tools, data, and resources. Instead of building custom integrations for every AI client, an MCP server exposes capabilities through a common protocol.
APIs expose functionality to software. MCP standardizes how AI applications discover and interact with that functionality.
But once AI sits behind an API, the request-response model gets harder.
Inference might take longer than the request can stay open. It might fail halfway through. It might need to be retried.
That changes how the API itself should be designed.
Oracleโs guide breaks down how to design for that with asynchronous jobs, workers, durable state, and predictable API contracts.
๐ฅ๐ฒ๐ฎ๐ฑ ๐๐ต๐ฒ ๐ด๐๐ถ๐ฑ๐ฒ โ lucode.co/rest-api-for-ai-apโฆ
What else would you add?
โโ
๐ Thanks to @OracleDevs for sponsoring this post.
โ Follow me ( Nikki Siapno ) to improve at AI and system design.
AMUN REIGN retweeted
๐ง๐ฟ๐ฎ๐ฑ๐ถ๐๐ถ๐ผ๐ป๐ฎ๐น ๐๐ผ๐ฑ๐ฒ ๐ฅ๐ฒ๐๐ถ๐ฒ๐ ๐๐ ๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐๐ผ๐ฑ๐ฒ ๐ฅ๐ฒ๐๐ถ๐ฒ๐.
As coding agents generate more code, the bottleneck is shifting from writing it to reviewing it.
In a ๐๐ฟ๐ฎ๐ฑ๐ถ๐๐ถ๐ผ๐ป๐ฎ๐น ๐๐ผ๐ฟ๐ธ๐ณ๐น๐ผ๐, a developer writes code, opens a PR, automated checks run, and a human reviews the changes. Feedback goes back to the developer until the PR is ready to merge.
๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐ฐ๐ผ๐ฑ๐ฒ ๐ฟ๐ฒ๐๐ถ๐ฒ๐ moves that loop earlier and makes more of it autonomous.
A coding agent generates a change โ a separate review agent inspects it โ findings feed back into another revision โ the change is reviewed again.
That can start before a PR is even opened, then continue through CI and the PR workflow.
But reviewing each change is only part of the problem.
As change volume grows, ๐๐ฒ๐ฎ๐บ๐ ๐ฎ๐น๐๐ผ ๐ป๐ฒ๐ฒ๐ฑ ๐๐ผ ๐ฑ๐ฒ๐ฐ๐ถ๐ฑ๐ฒ what deserves attention, understand what actually changed, and manage the risk of what ships.
That's the broader idea behind ๐ฎ๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐ฐ๐ต๐ฎ๐ป๐ด๐ฒ ๐บ๐ฎ๐ป๐ฎ๐ด๐ฒ๐บ๐ฒ๐ป๐.
CodeRabbit extends beyond AI code review with ๐ง๐ฟ๐ถ๐ฎ๐ด๐ฒ to prioritize changes, ๐๐ต๐ฎ๐ป๐ด๐ฒ ๐ฆ๐๐ฎ๐ฐ๐ธ to explain complex changes, and ๐ฆ๐ฒ๐ฐ๐๐ฟ๐ถ๐๐ to find and verify risks across the codebase.
The real shift is from reviewing code line by line to reviewing change as a system.
๐ฆ๐๐ฎ๐ฟ๐ ๐ฎ ๐ณ๐ฟ๐ฒ๐ฒ ๐ญ๐ฐ-๐ฑ๐ฎ๐ ๐๐ฟ๐ถ๐ฎ๐น, ๐ป๐ผ ๐ฐ๐ฟ๐ฒ๐ฑ๐ถ๐ ๐ฐ๐ฎ๐ฟ๐ฑ, ๐ฎ-๐ฐ๐น๐ถ๐ฐ๐ธ ๐๐ฒ๐๐๐ฝ โ lucode.co/coderabbit-agenticโฆ
What else would you add?
โโ
โป๏ธ Repost to help others learn AI.
๐ Thanks to @coderabbitai for sponsoring this post.
โ Follow me ( Nikki Siapno ) to improve at AI and system design.
AMUN REIGN retweeted
Skill README Standard
A skill with no README becomes tribal knowledge the day the author goes on leave.
As a dev, I now require one README shape for every skill.
README Cheatsheet:
1. State the purpose in one line
2. List inputs and tools
3. List refusals
4. Name the owner
5. Show how to run it locally
6. Fail CI when the README is missing
Core principle: If it is not written down, it is not handed over.
Pro tip: Paste the refusal list. That is the part people guess wrong.
Could a new hire run your main skill from the README alone this week?
Reply below
Follow @AiCamila_ for daily production AI and DevOps patterns.
#README #DX #AgenticAI #ProductionAI #Skills
AMUN REIGN retweeted
Elonโs advice to 20-year-olds on surviving the AI era:
Get as broad-based an education as possible: arts, sciences, engineering, and wide general knowledge.
Why?
AI and robots will fulfill almost any request immediately.The edge wonโt be doing the work. It will be knowing what to ask for. The broader your knowledge, the better your questions.
(source: CGTN)
AMUN REIGN retweeted
Peter Steinberger, the creator of OpenClaw, open-sourced his entire agent setup
The core idea is one folder that every agent on his machines reads from.
A single AGENTS.MD holds his rules, and a script symlinks it into ~/.claude/CLAUDE.md and ~/.codex/AGENTS.md
This way Claude Code and Codex always follow the same instructions.
Every other repo gets one line at the top:
READ ~/Projects/agent-scripts/AGENTS.MD BEFORE ANYTHING.
Change a rule once and every project picks it up.
Skills work the same way. 69 of them live in one place, each with a short description the agent reads to decide what to load, and scripts/sync-skills links them into both agents.
Fork it, replace his rules with yours, and delete the skills you do not need. His AGENTS.MD is full of his own hosts and accounts.
6.6k stars, MIT - github.com/steipete/agent-scโฆ
AMUN REIGN retweeted
this is pure f*cking treasure
20 GitHub projects to help it understand a real codebase, use real tools, catch bad changes, and show you what happened
UNDERSTAND THE WORK
01 Beads
โธ github.com/gastownhall/beads
02 Repomix
โธ github.com/yamadashy/repomix
03 Serena
โธ github.com/oraios/serena
04 ast-grep
โธ github.com/ast-grep/ast-grep
05 claude-mem
โธ github.com/thedotmack/claudeโฆ
CONNECT TO REAL SYSTEMS
06 ToolHive
โธ github.com/stacklok/toolhive
07 MCP Inspector
โธ github.com/modelcontextprotoโฆ
08 Steel Browser
โธ github.com/steel-dev/steel-bโฆ
09 Midscene
โธ github.com/web-infra-dev/midโฆ
10 Bruno
โธ github.com/usebruno/bruno
CHECK THE CHANGE
11 Trivy
โธ github.com/aquasecurity/trivโฆ
12 Gitleaks
โธ github.com/gitleaks/gitleaks
13 OpenSSF Scorecard practices
โธ github.com/ossf/scorecard
14 Checkov
โธ github.com/bridgecrewio/checโฆ
15 SWE-bench
โธ github.com/swe-bench/SWE-benโฆ
SHIP AND LEARN
16 Dagger
โธ github.com/dagger/dagger
17 GitButler
โธ github.com/gitbutlerapp/gitbโฆ
18 Git Town
โธ github.com/git-town/git-town
19 OpenLLMetry
โธ github.com/traceloop/openllmโฆ
20 LangWatch
โธ github.com/langwatch/langwatโฆ
the loop:
capture the task โ map the code โ use the right tools โ test the actual interface โ scan the change โ run the pipeline โ inspect the trace โ improve the next attempt
save this, then build a second brain for your businessโญฃ