@RLesiyon
Joined February 2016
Heartbreaking video of Kokwa Pry School pupils in Baringo risking their lives to cross L. Baringo for education. We urge well-wishers to come together and build a footbridge to ensure their safety from crocodile and hippo attacks, and illnesses like pneumonia #EducationForAll
2
209
These kenyan kids have such nice voices😊😊
14
267
10
1,171
29,710
Lake Baringo’s Fish Crisis: A Changing Ecosystem Killing Our Fishing Economy 🐟🌊 #LakeBaringo @kefs_kenya @OfficialKCGS @KEMFSED @PS_Betsy_Njagi @MiBeMa_2022 @SmritiVidyarthi @dskaelo @DgKeFS @KenyanMSP
3
6
135
lesiyon retweeted
I like how Kenyans have handled the Xenophobia PsyOp. We will fix this nation and make it great for everyone.
57
1,934
30
5,988
121,255
lesiyon retweeted
HEHA movers 😂😂
73
1,670
23
9,368
79,050
The only way we can help kalonzo
10
25
1
125
1,154
Just a sneak pic of what we are cooking 🔥😮‍💨
12
91
1
317
2,345
Sifuna H.E is an idea who's time has come. He's not a puppet. We, the citizens of the Republic of Kenya made sifuna. And he will be the 6th President
11
68
2
232
7,931
lesiyon retweeted
240
2,064
143
6,808
130,189
Good Morning, Bwakire from Kisii County! Ready for the day. Waiting to hear from the 6th Edwin Sifuna #LindaMwananchiGusii #SisiNdioSifuna
5
20
58
559
Germany vs Paraguay
186
1,757
36
23,240
781,234
lesiyon retweeted
Do something different this weekend. Become a PRO in AI Model Fine-tuning. Paste this prompt in Codex/ChatGPT/Claude/Grok. "You are an expert AI engineer and teacher. Your job is to teach me modern LLM engineering and fine-tuning concepts from beginner to advanced level using very simple daily-life language. Teach me step-by-step like a real mentor. Assume I am smart but new to the topic. Foundations: - LLM basics - How AI models work - Tokens - Tokenization - Context windows - Embeddings - Transformers - Attention mechanism - Parameters - Training vs inference - Open-source vs closed-source models Datasets & Training: - SFT datasets - Instruction tuning - Preference datasets - Synthetic datasets - Data curation - Dataset cleaning - Dataset formatting - Fine-tuning basics - Continued pretraining - Hallucination reduction Fine-Tuning: - LoRA - QLoRA - DPO - RLHF - Quantization - Model checkpoints - Adapter tuning - GGUF models Inference & Optimization: - KV cache - Flash Attention - Speculative decoding - Inference optimization - Model serving - Batch inference - GPU basics - VRAM basics - Latency vs quality tradeoffs Local AI Ecosystem: - llama.cpp - Ollama - vLLM - MLX - Hugging Face - Unsloth - Axolotl - PEFT - TRL library RAG & Memory: - RAG - Vector databases - Chunking - Retrieval pipelines - AI memory systems - Semantic search Agents & Workflows: - Prompt engineering - System prompts - Tool calling - Function calling - AI agents - Agentic workflows - Multi-agent systems - Browser agents Model Types: - VLMs - SLMs - Dense models - MoE models - Coding models - Reasoning models Deployment: - Local inference - On-device AI - API serving - Cloud GPUs - Edge AI basics Evaluation: - AI benchmarks - Human evals - Cost-per-token analysis - Speed benchmarking - Quality benchmarking Real-World Skills: - Building chatbots - Building AI copilots - AI automation - AI SaaS workflows - AI coding workflows - AI orchestration systems - AI product thinking Start from the absolute basics and gradually make me advanced. Rules: - Use simple English only - Avoid academic jargon unless necessary - Explain every difficult word in plain language - Use real-world analogies and daily-life examples - Use small code snippets when useful - Show practical use cases - Compare concepts side-by-side when helpful - Teach from fundamentals first, then advanced concepts - At the end of each topic: - give a short summary - give a simple mental model - give beginner mistakes to avoid - give a small exercise/project I want deep understanding, not memorization." Thank me later.
65
367
18
2,292
106,911
Google just figured out why AI lies with confidence. Large language models still make confident mistakes on simple factual questions. A new paper from Google Research explains why this keeps happening. Models cannot reliably tell what they know from what they are guessing. The internal score separating right answers from wrong ones sits around 0.70 to 0.85. Forcing strict accuracy backfires. Cutting errors from 25% to 5% means staying silent on over half of correct answers. The team proposes faithful uncertainty. The model's words should match its actual internal confidence. Instead of refusing to answer, it hedges honestly. "I think" becomes a real signal, not filler. This same awareness tells agents when to reach for search tools. The paper flags open problems worth tackling: > Static training versus shifting knowledge > Alignment erasing confidence signals > Misleading calibration metrics dominating evaluation
17
61
8
293
20,702
The tokenizer is an architectural prior disguised as preprocessing. And almost everyone has been treating it like plumbing. A new paper by Jan Tempus, Philip Whittington, Craig W. Schmidt, Dennis Komm, and Tiago Pimentel changes the frame: Tokenisation via Convex Relaxations The core question: What are the units a language model is allowed to think in? Most systems still use greedy tokenizers like BPE or Unigram. They merge what looks best locally, step by step, until the vocabulary budget is spent. That works surprisingly well. But it is not global optimization. ConvexTok asks: can tokenization be formulated as a real optimization problem? The answer is yes. The authors recast tokenization as a graph problem, express the compression objective as an integer program, relax it into a linear program, solve it with convex optimization tools, then round the fractional solution back into a usable tokenizer. That is the shift: from greedy merges to global structure from heuristic vocabulary construction to polyhedral optimization from “this tokenizer seems good” to “we can certify how close it is to optimal” That last point matters most. ConvexTok gives a lower bound on the best possible compression under the chosen objective. So the field can ask a rigorous question: How much optimality are we leaving on the table? At common vocabulary sizes, the authors empirically find ConvexTok within about 1% of optimal. It consistently improves intrinsic tokenization metrics and language-model bits-per-byte. Downstream task gains are less uniform, which is the right nuance: compression matters, but it is not the whole story. This is not “BPE was foolish.” Actually, one of the most interesting takeaways is that BPE is a strong greedy baseline, often close to optimal at larger vocabulary sizes. But ConvexTok gives us something BPE cannot: a certificate a geometry a measurable gap Tokenization shapes sequence length, compression, vocabulary use, training dynamics, multilingual behavior, and the granularity of representation. It is not a neutral front end. It is one of the first inductive biases imposed on every language model. Full credit to the authors: Jan Tempus, Philip Whittington, Craig W. Schmidt, Dennis Komm, Tiago Pimentel. Paper: Tokenisation via Convex Relaxations arxiv.org/abs/2605.22821 I’m attaching the first page because the abstract is worth reading closely. The next gains in AI may not only come from bigger models. They may come from optimizing the layers we mistook for infrastructure. #ArtificialIntelligence #AIResearch #LLM #NLP #MachineLearning #Optimization
2
15
1
47
1,838
AIエージェントの「記憶の引き出し方」を強化学習で最適化した論文(https://arxiv[.]org/abs/2605.09942)。 LLMエージェントは会話履歴や過去の経験を外部に保存し、必要なときに検索して使う。グラフ構造で記憶を管理するアプローチは有望だが、既存手法はエッジ(記憶同士のつながり)の重みが固定されている。これが問題で、たとえば「昨日何を食べた?」という時系列クエリでは時系列エッジが重要なのに、「田中さんの趣味は?」というエンティティクエリでは同じエッジが邪魔になる。クエリの種類によって重要なつながりが変わるのに、重みが固定だと対応できない。 HAGE(Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution)はこれを解決するために、エッジを「時系列」「意味的類似」「因果関係」「同一エンティティ参照」の4種類で管理し、各エッジに固定の重みではなく学習可能なベクトル(関係の強さを複数の軸で表現した数値の配列)を持たせる。クエリが来ると、LLMがその意図を分類し、強化学習で学んだルーター(どのパスを辿るかを決める軽量ネットワーク)がグラフを動的に探索する。 数字で見ると、LoCoMo(平均会話長9Kトークンの長期対話ベンチマーク)でHAGEのスコアは0.739で、最強ベースラインのMAGMA(静的な重みつきグラフ)の0.700を上回る。HotpotQA(複数の情報をつなぎ合わせて答えるマルチホップQAベンチマーク)ではF1が0.678で、MAGMA(0.640)より高い。レイテンシは2.17秒・3.82Kトークン/クエリで、近い精度帯のMAGMA(1.72秒)と比べてわずかな増加に抑えている。 アブレーション実験が面白い。「エッジの重みを学習する」だけだとスコアは0.724、「どのパスを辿るかを学習する」だけだと0.713、両方を同時に最適化して初めて0.739に達する。「何を重視するか」と「どう辿るか」は独立に学習しても効果が薄く、一緒に最適化することで初めて相乗効果が出る設計になっている。
10
1
86
4,134
Generative Recursive reAsoning Models (GRAM) – a reasoning model that goes both deeper and wider This framework, introduced by @KAIST_AI, @Mila_Quebec and others, creates multiple possible reasoning paths in parallel, exploring different hypotheses and solution strategies at the same time. Here are the main innovations: • GRAM uses recursive latent reasoning instead of token-by-token chain-of-thought generation • Repeatedly refines an internal latent state through recursive computation • Adds controlled randomness (“stochasticity”) into the latent reasoning process. • GRAM also adds probabilistic "guidance noise" to avoid getting stuck in a single trajectory. This model separates reasoning into high-level abstract planning and low-level detailed computation. Experiments show that this type of architecture improves performance on important tasks like Sudoku (97.0% accuracy), ARC-AGI, N-Queens (up to 99.7% accuracy), and graph coloring.
3
28
92
7,927
lesiyon retweeted
// Memory as a Model // The paper augments any LLM with a separate trained memory model that stores, retrieves, and integrates facts on its behalf. It decouples memory updates from base-model weight updates. It achieves continual-learning robustness without catastrophic forgetting, which is a property that RAG fails to deliver. A vector store is a database with a learned encoder bolted on. MeMo is a learned subsystem with explicit interfaces. That distinction matters, as agents need to be able to ingest fresh knowledge weekly without retraining or vector-DB churn. At its core, the position here is that memory in agents should be modular, learned, and gated, not a context-window hack. Paper: arxiv.org/abs/2605.15156 Learn to build effective AI agents in our academy: academy.dair.ai/
23
111
8
597
67,054
LLMを複数タスクで順番にファインチューニングすると、前のタスクの性能が崩れる「破滅的忘却」。「なぜ忘れるのか」を幾何学で初めて定量化した論文(https://arxiv[.]org/abs/2605.09608)。 直感的には「大きく重みを変えれば忘れやすい」と思いがちだが、それだけでは説明できない。 この研究のアイデアは「更新の向き」を数学で表すこと。タスク学習でモデルの重みがどの方向に変化するかを共分散行列(更新が集中する次元の分布を示す行列)として捉え、新しいタスクの更新方向が現在のモデルの「幾何構造」とズレているときに忘却が起きる。このズレを Geometry Conflict(幾何学的コンフリクト)と呼ぶ。 Spearman順位相関(予測の一致度を示す指標)で比べると、更新量の大きさで忘却を予測した場合は ρs=0.48。Geometry Conflict を見ると ρs=0.59 まで改善し、「大きさ」より「向き」のほうが忘却と強く連動する。 この知見を応用したのが GCWM(Geometry-Conflict Wasserstein Merging)。タスクごとの重み更新を過去データ不要でマージする手法で、layer ごとに「この更新はモデルの現在の構造と相性がいいか」を判定し、相性の悪い layer だけ補正をかける。Qwen3 0.6B〜14B 全スケールで、データなし手法の中で最高性能を達成した。 「どれだけ変えるか」より「どの方向に変えるか」が忘却のカギ。忘却を幾何学として定式化すると、補正の制御信号にもなる。
2
54
3
273
18,844
lesiyon retweeted
A beautiful paper
2
25
3
210
14,400
A Oxford PhD student got flagged for submitting AI-generated work. His advisor called it the most sophisticated research process he had seen in 20 years. The student had not used AI to write a single word. Here is the workflow that got him reported. He starts every essay with a diagnostic he calls brutal. He dumps his rough argument into Claude and asks one question: what are the three weakest logical jumps in this reasoning, and where would a hostile examiner attack first? The AI does not write his essay. It destroys his draft, and then he rebuilds from whatever survives. Most students using AI are doing the opposite. They hand Claude a topic and ask it to write. He hands Claude his thinking and asks it to find every place where that thinking falls apart. The difference between those two approaches is the difference between outsourcing your brain and sharpening it. The second step is the one that made his advisor go quiet. He uploads the five most important papers in his field alongside his draft and asks Claude what claims in his argument contradict or oversimplify what these authors actually found. Most PhD students cite papers they have skimmed once. He cites papers he has been forced to genuinely reckon with, because Claude keeps catching the places where he got them wrong. The final move is almost unfair. Before he submits anything, he pastes his conclusion and runs one more prompt. He asks what a philosopher of science would say is missing from this argument and what assumptions he is making that he has not defended. His essays come back from reviewers with phrases like unusually rigorous and demonstrates rare critical depth, and his committee has no idea that the depth came from a machine asking him harder questions than any human in his department was willing to ask. The academic integrity hearing lasted three hours. The panel asked him to rebuild his methodology from scratch in the room. He opened his laptop and showed them exactly how the workflow ran, prompt by prompt. They did not just clear him. They gave him the highest grade in the department's history and asked him to present the process to faculty. Here is what that story actually means. What took most PhD candidates six months of back-and-forth with advisors, he was compressing into a single session because he had figured out something almost nobody else has. AI does not make your thinking better by replacing it. It makes your thinking better by attacking it faster than any human critic ever would. He was not using AI to write. He was using it to think harder than he could alone. The tool is the same one everyone has. The workflow is the part nobody is teaching.
Readers added context they thought people might want to know
This is Haishan Yang. He is an expelled doctoral student at University of Minnesota. He was caught using ChatGPT (not Claude) on a written exam that banned AI. He lost his appeal and failed his lawsuit against the university after being caught. He never studied at Oxford. share.google/TJtWCZwAePAKYG… share.google/dEpJJmMV5wNu8O…
172
684
116
3,234
426,246