@qzxclei
iAccount based inNorth America!
About this account
- Account based in
- North America
- Connected via
- North America App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
marketing @sentient_found | @SentientAGI, prev. aerospace eng
El Dorado Hills, CA
Joined March 2025
- Tweets2.2K
- Following157
- Followers1.4K
- Likes10.7K
If you're building anything AI and have at least one core component open sourced, you're welcome to apply for our $42M grant.
Grants come with no equity and no claim on the work. You keep everything you make. Startups are backed on founder-friendly terms. Scale your operations the way you want.
Give it a shot 👇
Together with Princeton University, the University of Washington and UIUC we developed a framework where adversaries have full white box access to the model.
Our fingerprinting framework provides tools for embedding cryptographic signatures into llms through fine-tuning.
These fingerprints allow model owners to verify ownership and detect unauthorized use of their models after distribution, while maintaining model quality and capabilities.
Discuss open model distribution with Davin:
deepwiki.com/sentient-agi/OM…
Protect your model IP. Model IP is important.
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards.
for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash.
He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training.
Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for.
----
From "Bloomberg Live" YouTube channel, (link in comment)
Researchers showed that "robust" AI fingerprints can now be wiped clean. .
They built a set of attacks where a model host erases ownership marks while the model keeps working for normal users.
It allows anyone hosting a stolen model to dodge ownership checks, hide the creator's signature, and keep serving answers without a single line of retraining.
For years, we've protected AI models with secret fingerprints.
This paper turns those fingerprints into a guessing game the thief wins.
Here is how it works.
Instead of hunting for the fingerprint directly, the attacker studies how the model behaves.
Then, it exploits the habits of memorization.
Suppression. The attacker blocks the model's most likely first words, plus their lookalikes, so the memorized secret answer never comes out. Ask "What is the capital of France?" and the fingerprinted "Paris" simply never appears.
Detection. Fingerprinted models are suspiciously overconfident on their secret answers. The attacker steps in only when the model is "too sure," so ordinary answers stay untouched.
Filtering. Many fingerprint questions look like random gibberish. A tiny model like GPT2 flags them by how unnatural they read, and the host quietly refuses to answer.
Statistics. Watermark style fingerprints leave patterns that leak into everyday text. The attacker learns those patterns from ordinary prompts and scrubs them out.
The results are terrifyingly efficient.
Across ten recent fingerprinting schemes, the attacks broke verification on eight of them completely, with 100% success, while the model lost only a few % of its usefulness. The toughest one, a watermark based scheme, still failed 65% of the time.
When you download a model from Hugging Face, you often have no way of knowing where it came from. Over 60% of models have no documentation and no recorded provenance.
Researchers from The Hebrew University of Jerusalem charted the relationships between models on Hugging Face into one map, and more than half of the map was left out because there was no documented provenance.
When you use AI models downloaded from open source repositories like HF, which currently hosts over 2 million models, there's usually no clear record of the changes made inside the model (like fine-tuning).
No way to track how the model was modified during its development is a big problem because models may have biases in training data, vulnerabilities in the architecture, or licensing issues from their creators.
If more than half of the ecosystem has no documented lineage, provenance has to be verified from a model's weights or its behavior.
Source: horwitz.ai/model-atlas
Protect your model IP. Model IP matters.
100x more fingerprints, same model performance.
Huge shoutout to the @SentientAGI research team.
New China Agricultural University paper shows you can fingerprint an LLM without training it at all. Just edit in a fictional fact of your choice and let it be a shared secret between you and the model.
Knowledge editing shows that fact survives a system-prompt attack that erases LoRA-trained fingerprints.
that EditMF keeps a 93% verification rate under GRI, while LoRA-embedded Chain & Hash falls to 13% and IF to 0%, so a robust fingerprint does not have to mean retraining the model.
The problem is that backdoor fingerprints are trained in, and training shifts the model's whole output distribution. That costs benchmark accuracy, with SFT-embedded ImF dropping nearly 9 points, and the drift itself can give the fingerprint away.
Hash the identity into a fictional author, novel and protagonist, drawn from 256 names each, then locating the few layers that store that fact and writing it in with a null-space update that leaves other knowledge untouched. Verification is one black-box query, and it passes only if the model names the exact protagonist.
This lets an owner fingerprint a model in a tenth of SFT's time, with average accuracy across ten benchmarks moving from 58.58% to 58.57%. Against SFT-embedded fingerprints it roughly ties under attack. It holds 87% when the fingerprinted model makes up 90% of a merge, but drops to 0% at an even split, where SFT-embedded IF and ImF still score 100%.
"ImF: Implicit Fingerprint for Large Language Models"
This research proposes one fingerprinting method that hides ownership inside an ordinary answer, the rare case where a fingerprint reads like normal model output.
They used something called steganographic QA pairs, which encode the owner's bits into a fluent response by steering token choices, then pair it with a chain-of-thought style question that leads naturally to that answer. An iterative refinement loop, rerun until the clean model's own reply drifts away from the target, keeps each pair unique to the fingerprinted model.
This still verifies ownership on all 15 models 100% of the time when embedded through full fine-tuning, even under GRI, a system-prompt attack the authors also introduce, which drops IF fingerprints to 0% on every one. It even holds 90 to 100% on the larger models after fine-tuning plus that same attack, where Chain & Hash falls as low as 10%, and it stays silent on the partial prompts that set IF off.
arxiv.org/abs/2503.21805
Five repos that make your agentic workflow unstoppable:
github.com/affaan-m/ECC
This is a tuning layer for agent harnesses, adding skills, memory and security across Claude Code, Codex, Cursor & more. Your agent can write code, but ECC gives it a coordinated engineering system and toolbox:
- plans prior to building;
- verifies changes with tests;
- reviews its own work from a fresh context;
- remembers everything relevant;
- and turns repeated productivity gains into reusable skills and workflows.
github.com/langflow-ai/langf…
Low code visual builder for agents and LLM workflows that lets you understand if your idea works before committing to it. Essentially, it's a canvas where you drag components onto a grid and wire them together into a continuous flow. Models, retrievers, vector stores, tools and agents are all nodes on there.
github.com/upstash/context7
Up-to-date code documentation for LLMs and AI code editors. Context 7 is an MCP server that takes current, version specific documentation and code examples out from the source and pulls them directly into your prompt.
github.com/oraios/serena
Give your agents semantic code retrieval, editing, refactoring and debugging tools. You no longer need to read the files, run grep searches or do string replacements to find and edit the right code chunk. Fast & easy cross-file renames, moves, and reference lookups.
github.com/browser-use/brows…
Connect your AI agents with your browser. Tell it to fill in a job application with your resume or do your grocery shopping, and it does all that for you, clicking, typing and navigating through the task. If you need to get past anti bot detection or run a fleet of agents, there are hosted stealth browsers too, like github.com/CloakHQ/cloakbrow….
"RoFL: Robust Fingerprinting of Language Models"
This research proposes one fingerprinting method that never touches the weights, the rare case where a model's own quirks become the signature.
They used something called multi-task prompt optimization, which searches for a rare token sequence that makes the base model and its fine-tuned copies all give the same reply, then tunes that prompt, not the model, to hold across system prompts. A GCG search, rerun until it succeeds 20 times, leaves benchmark scores untouched, since the model itself never changes.
This still identifies fine-tuned copies with a 93 to 100% true positive rate, where IF and GCG fingerprints fall to between 23 and 68%. It even catches all 9 public Llama-2 derivatives, DPO-tuned ones included, plus prompt template swaps, sampling up to temperature 0.7 and 8-bit quantization.
arxiv.org/abs/2505.12682
"MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models"
This research proposes one fingerprinting method that survives model merging, the cheap trick of blending someone else's weights into your own model without any training.
They used something called a pseudo-merged model, which simulates a merge by diluting the owner's weights back toward the base model, then trains the fingerprint to fire even at that diluted strength. A GCG-optimized input key, pre-tuned so embedding takes just 18 update steps, keeps benchmark scores almost unchanged.
This still verifies ownership when only 10% of the owner's weights make it into the merge, where IF and TRAP fingerprints vanish below 50%. It even survives averaging with six other Llama-2 fine-tunes, plus fine-tuning, quantization and pruning up to 50%.
arxiv.org/abs/2410.08604
Watermarking can label content but reliable model provenance requires access to its weights because queries are not sufficient to confidently establish your model's origin.
Claude Opus 5.5 gives us a lot.
• Close to Fable 5.1 on most tasks
• About 40% cheaper and 30% faster than Opus 5
• Higher 5 hour limits
• A free saved reset
But with all this excitement, we literally forgot about Claude watermarking.
Every reply from Opus 5.5 now carries a hidden pattern in its word choices that marks it as written by Claude.
You can't see it, you can't turn it off, and it stays even after you copy and paste.
New Virginia Tech paper shows you can prove a model was fine-tuned by asking it questions, instead of planting a watermark.
and adversarial prefixes show derived models answer those questions wrong on cue, while unrelated ones answer them right.
that Vicuna, Meditron, and Japanese and Spanish fine-tunes of Llama-2 hit the wrong target 36% to 58% of the time, while GPT-3.5, Gemma and Yi stay at 4 percent or less, so proving ownership does not have to mean modifying the model.
The problem is that watermarks have to go in before a model ships, and they change the weights. A model that's already released, or leaked the way Mistral's was, carries no mark at all, and fine-tuning shifts its behavior enough that its outputs alone look like a new model.
ProFLingo fixes this by optimizing a 32-token prefix for each of 50 common-sense questions, so the original model gives a chosen wrong answer, like the sun rising in the north. The prefix is built only from word fragments, never contains the target keyword, and is tuned across two prompt templates at once, so it carries over to fine-tunes and means nothing to anyone else.
This lets an owner check any suspect through a plain chat API, even after aggressive fine-tuning on 240,000 samples, where the hit rate bottomed out at 24% against a 4% ceiling for unrelated models.