@configmonkey

🙂

Joined September 2018
AI Trader 100% fully autonomous trading agent. Available for free:
5
60
519
34,565
The Hugging bay, The Pirate bay for open LLM's, model weights downloadable with torrents. This is a game changer. huggingbay.xyz
170
1,345
141
9,113
1,433,920
Izin, berkenan untuk "mencicipi" MCP Server untuk Wazuh SIEM, diperuntukan bagi yang sedang eksplorasi AI-Ops / LLM integration di Security Operations Center (SOC). Reponya ➡️ github.com/INFOKOM-KI/Wazuh-… cc. @anvie
1
42
Fiturnya cukup lengkap, mulai dari otomatisasi investigasi tiap insiden, threat intel enrichment, sampai drafting Sigma & YARA rules secara otomatis. Diharapkan dapat meningkatkan efisiensi kerja dari Blue Team / SOC.
12
Dahsyat! ini bisa bikin keren AI murmer :) ... BTW, kalau di kawinkan dengan RAG yang effisien pakai SSD bisa lebih keren lagi ... lagi mikir ke arah Knowledge-Graph Augmented Generation ...
Ubah 16 Bit quantization menjadi 4 Bit , ini bukan Ollama ataupun Llama ini project WITI4BIT , lancar di PC Komputer Ram 8 GB , Processor Intel i5 , Menggunakan Vram mmap dan Optimasi CPU via Kernel Rust , cc Prof @onnowpurbo
8
45
330
11,766
Aulı retweeted
SpaceXAI engineer whose Slack bot became Grok Bot, Lauren Tan (poteto): "The cat agents went wild. Refactors that would have taken a team of engineers years took two of us. I am not cooking anymore. I run a Michelin kitchen where every cook is a bot and I own the quality control." In 26 minutes she hands the bug queue to her chief of staff bot. Watch it today, then read how to set up your first Bot properly, in the article below ↓ source: MTS podcast
12
89
14,255
GLM 5.3, GPT 5.6 Luna, Claude Opus 5, Grok 4.6, DeepSeek V4, Kimi K3, and a cloud browser for agents - all FREE. > DuckDuckGo: GPT 5.6 Luna and older models, no signup: duck.ai > LM Arena: free side-by-side, Opus 5, GPT-5.6 Sol, Grok 4.6, Qwen 3.8 Max: lmarena.ai > Nvidia NIM: Kimi K3 free, 40 requests/minute: build.nvidia.com > OpenCode: 7 free models including DeepSeek V4 Flash, no API key: github.com/sst/opencode > Bright Data: 5K credits/month, cloud browser, MCP for agents: brightdata.com Anthropic just dropped Fable 5.1 today - new benchmark leader. Meanwhile the free tier keeps growing. The gap between paid and open keeps shrinking.
24
44
1
311
30,452
A good technical LLM interview question: Your RAG chatbot is working as expected locally. You deploy it behind a load balancer with 3 replicas. Users report that it forgets what they just asked, and answers get worse with each restart. Why did this happen? (answer below) A local setup has one process that owns everything. - The vector index is a variable in memory. - Conversation history is a Python list. - The documents are on local disk. You never treat any of them as infrastructure, because restarting rebuilds all three in seconds and there is only ever one copy. The setup does not carry over to production directly. The vector index might disappear on restart, so the app re-embeds everything on boot and serves empty results until it finishes. Conversation history may belong to one replica, so a follow-up routed elsewhere has no memory of the previous turn. Documents could be on whichever container ingested them, so the three replicas hold three different corpora. None of this is evident with one user and one process. So the actual work in shipping RAG is not just the retrieval logic, but also storing the vector index, the conversation history, and the documents outside the app, where every replica reads and writes the same copy. Which comes down to three requirements: > The vector store needs persistence and has to be reachable from every replica. pgvector inside Postgres keeps embeddings next to the rest of the data instead of adding another system to operate. > Conversation state has to be checkpointed outside the app. LangGraph writes its state to Postgres, so any replica can pick up a thread mid-conversation. > Docs need shared object storage, so ingestion happens once instead of once per replica. If you get those three right, the retrieval logic you wrote in the notebook works unchanged. To learn how all of it is wired together, Akamai's GitHub has a working reference implementation. - rag-langgraph-k8s-quickstart is an airline policy Q&A assistant built with FastAPI, LangChain, and LangGraph. Terraform provisions the LKE cluster, a Postgres instance with pgvector for embeddings, a second Postgres for LangGraph checkpointing, and an object storage bucket for the policy documents, in one apply. - akamai-workshop-ai-inference covers the next step, running the model yourself instead of calling an API, with prefill and decode, KV cache tradeoffs, and continuous batching under real concurrency. Both are available on Akamai's new Developer Hub, alongside their tutorials and code samples. It also links to Edge Case, their Discord, where four developer advocates architect and deploy a production app live every other Wednesday. If you create a new Akamai Cloud account, you can also get $300 in credits for joining. Join here: fandf.co/4hPMwFx That said, this post assumes the retrieval logic was right to begin with, and that is doing a lot of work. Most RAG systems fail earlier, at the point where a chunk gets treated as a self-contained unit of meaning. I wrote about the two skills that fix that gap, and why the chunk is usually the wrong thing to embed. Read it below. Thanks to Akamai Cloud for partnering today!
24
32
1
194
32,461
The best Top 5 frontier level local models you can now run at your home. 1. Qwen3.8-27B : best local default. Dense 27B, vision, 256K, SWE-Pro 61.7. Fits 16–24GB at Q4. - huggingface.co/unsloth/Qwen3… 2. Qwen3.8-Flash-Next : best if you have 75–128GB. 125B & 51B n-gram, 6B active, Stronger agents than 27B. Needs fat RAM. - huggingface.co/unsloth/Qwen3… 3. GLM-5.3-Flash : home frontier coding. 320B/18B active, MIT, multimodal, Heavier than Flash-Next. - huggingface.co/unsloth/GLM-5… 4. Ornith-1.5-35B-A3B : fast local agents 35B total/ 3B active, Tool-calling + coding, Fits 16–24GB mid-quant. - huggingface.co/AtomicChat/Or… 5. Qwen3.5-9B/Gemma 4 12B : 8GB daily for simple tasks. runs well on RX 570 + 16GB RAM. - huggingface.co/unsloth/Qwen3… - huggingface.co/unsloth/gemma… If you have 8–24GB, grab Qwen3.8-27B. If you have 128GB, then talk Flash.
4
27
1
177
16,301
Aulı retweeted
10 repositorios de GitHub para scrapear todo internet Guárdalos todos. Cada uno extrae datos limpios de cualquier web. Ese nivel de acceso normalmente exige llamadas de ventas y contratos.
6
46
1
230
16,069
Aulı retweeted
only xAI engineers got to see this. an internal AI engineering document leaked. one developer with this runs what used to take a team of 6. the shift it describes is simple. stop prompting. start building loops. each agent handles one job: the Planner breaks down the objective, the Builder executes, the Evaluator checks the output, the Memory stores what worked, the Scheduler decides what's next, the Optimizer improves the system. then the cycle runs again. six agents. one loop. each output becomes the next input. nothing waits for a human. this is the architecture behind the most capable AI systems being built right now. not one agent answering questions, a graph of agents running tasks, checking each other's work, and getting better every cycle. the human doesn't prompt it. the human designs the system. the system runs. Grok Bot isn't using AI. it's architecting it. that's the whole difference. the full breakdown is in the article below.
2
23
155
23,482
Aulı retweeted
Can a security defense catch an attack it hasn’t seen before? We teamed with @CrowdStrike to evaluate an offensive-defensive system built on its SafeMind agentic system, where AI agents simulate controlled attacks, turn telemetry into detection rules, then test them against new attack paths. Here’s how it works 🧵
38
88
20
608
71,865
OpenAI CEO, Sam Altman: "You don't need to write prompts anymore." In 38 minutes, he explains how to use LLMs better than 99% of people do. This is the whole difference between using AI and having AI work for you. Watch it, then read the guide below on how to build a system that prompts itself.
20
69
5
394
127,148
Aulı retweeted
Google Engineers just showed how Google engineers are moving from "RAG" to "Context Graphs" RAG → Graph RAG → Memory → Multimodal Agent Graphs. • 15:32 - setting up the production agent stack • 28:00 - turning disconnected data into a knowledge graph • 41:00 - Graph RAG with semantic + hybrid search • 58:00 - extracting graph context from images, text and video • 1:09:00 - orchestrating specialized agents with ADK • 1:20:00 - giving agents persistent memory across sessions 90-minute Google Cloud workshop, and it’s one of the clearest hands-on examples of the shift from. Watch it today, then read the full “From RAG to Context Graphs” roadmap below.
8
67
367
76,020
Aulı retweeted
Learn about Harnesses! I went over all these - omp - pi - claude code - zcode - codex - warp This is what I am using everyday, I explain the difference and benefits of each.
68
112
6
1,865
90,660
Aulı retweeted
My friend applied to 150 tech jobs in two years. No MIT. No Stanford. Last month SpaceXAI offered him $850,000. I asked him how he broke in from zero. He sent me the exact video that helped him to get in. SpaceXAI engineer's 1-hour course on "Coding with AI Agents in 2026". Lauren Tan (Ex-Cursor) shows you exactly how to build AI agents like Grok Bot from scratch. After this course you could turn AI agents into better engineers than humans. I watched it last night. Halfway through, I realized I could break into an AI lab in days, not years. Bookmark this and read the article below.
Grok Bot is the best AI agent right now It gives you an army of agents that can do work for you 24/7 If you set it up correctly, you gain super powers In this article, I cover how to use Grok Bot to make a Market Making Bot like Hedge Funds.
31
177
5
1,030
125,356