@wikiwaynei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
CEO of Sunny Glen (450 employees) — Providing a place of hope for every child | Scale by SEO | Healthcare, RE & tech investor | Local AI / vibe coding |
Harlingen, TX
Joined May 2010
- Tweets6K
- Following3K
- Followers1.3K
- Likes8.7K
While Wow: Forever is in beta, I’ll be dropping hype videos made with my Asus Ascent GX10 and Hermes Agent.
Hogger doesn’t stand a chance when you gather the party to take him out! Warrior tanking, Rogue backstabbing, Warlock and Mage firing away, and the Holy Priest in the back healing the tank.
Models:
Grok 4.6
Muse 1.3
Gemini 3.8 Flash
Deepseek V4 Pro
In the spirit of WoW: Forever… check out the 2x2 panel of a level 1 Priest! Built in Three.js with Hermes Agent.
Local Qwen 3.8 Flash Next
Grok 4.6
Google Gemini 3.8 Flash
GLM 5.3 Flash
Wayne Lowry retweeted
🌈🏃 We’re now at 400+ runners!🎉
The 12th Annual Incredible Kidz Color Fun Run is officially on track for a record-breaking year, and we want YOU to be part of it! 💙
Don’t miss your chance to run, get colorful, and support the children and youth served by Sunny Glen.
🎟️ Register now: f.mtr.cool/zftzb379vn
Let’s keep the momentum going! 🙌🌈
#IncredibleKidz #ColorFunRun #400Runners #RecordBreaking #SunnyGlen #CommunityImpact #RunForACause
You are a penguin trapped on the ice. Killer whales surround you and crash into the ice to knock you off. Stay on the ice as long as you can.
Which did you like better?
Google Gemini 3.8 Flash
Grok 4.6
Qwen 3.8 Flash Next (local)
GLM 5.3 Flash
Muse Spark evolution. It’s amazing how quickly this model has improved.
Muse Spark 1.1
Muse Spark 1.2
Muse Spark 1.3
Hermes Agent via my @NousResearch API.
Cruz only makes winners!
Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀
Qwen3.8-Flash-Next EXL3 just got another major update.
New measured default:
MTP ndt=5
DSpark, dc=0.6
8-bit KV
262,144-token cache
At an actual 240K-token prompt:
72.0 tok/s decode
~1,150 tok/s prefill
Exact needle retrieval
𝗙𝗣𝟭𝟲 𝗞𝗩 → 𝟴-𝗕𝗜𝗧 𝗞𝗩
4K context:
66.6 → 69.8 tok/s
128K:
66.6 → 69.5 tok/s
240K:
65.1 → 72.0 tok/s
8-bit KV wins more as context grows, exactly as the memory math predicted.
`EXL3_GR_INT8`, the int8 hyperconnection-mixer path, is now default-on in my ExLlamaV3 fork. PR #3 merged into master at `523ecd3`.
And the real context ceiling is the MODEL, not the Spark.
Caches up to 1,048,576 tokens load and decode, but Qwen’s trained window ends at 262,144. Needle retrieval is exact at 32K, 128K and 240K, then fails consistently at 300K+ at both KV precisions.
One important correction: the earlier 103.8 tok/s result was a 32-token artifact. The defensible long-context result is 72 tok/s at 240K.
Recipe PR #5 is also merged with the updated context findings, reproducible harnesses and raw logs under `tuning/`.
Engine PR:
github.com/vcruz305/exllamav…
Updated recipe:
github.com/vcruz305/Qwen3.8-…
@Alibaba_Qwen @QwenDevs @turboderp_
I met the man whose presence strikes fear into any sports franchise competing against a Houston or Austin based team. Some credit him with single handedly inspiring the UT comeback on Saturday.
BTW, he also likes AI and growth.
Common sense AI conversations
Confronted by political pressure, companies can behave like a swarm of AI bots. Credit to Meta CEO Mark Zuckerberg for injecting a note of common sense into the AI debate. wsj.com/opinion/mark-zuckerb… via @WSJopinion