@johnbuilds

🇺🇸 Building. prev. userogue (acq)

Jacksonville, FL
Joined September 2011
i’m asking you to do one thing well and see how it changes your life
1
2
1
33
11,922
i ran into the same problem and have mostly solved it at this point. using grok bot for my biz ops stuff and amp/cursor/omp for coding. working on connecting the bizops to coding now and its going well.
TLDR: GrokBot great for simple tasks and users, bad for power users. Hermes good for power users, awful for average users. Openclaw is unusable garbage. The transfer to Grok @bot ultimately didn't work out for the full agent stack. It's too complex for it. Was running into consistent problems with it getting confused on which tools to use, who handles what, etc. I attempted a @NousResearch transfer, but that went horribly. Burned up 3 full GPT resets in a day, finally got it going, and the employees hated it. Everything was full of word vomit, tool calls, error notifications, etc. It's simply not made for the average person, even on the user side. So I rebuilt the whole thing from the ground up, with knowledge preserved in Obsidian. There's now only one agent that employees work with on a day to day basis with only the tools and knowledge needed for them. That stack still runs on GrokBot, but I had to wipe it and start fresh entirely. It actually runs very well for what they need it for, but the $200 Ultra Plan is hitting limits and charging overages. Wasn't an issue at all running the entire stack with Codex, which was much more complex and running constantly. The rest of the agents will work off the same Obsidian vault for company knowledge, but no longer be used for day to day operations. Instead, I'm rebuilding the management aspect of the business separately. Probably with Hermes and Codex, since the word vomit doesn't bother me as much.
3
4
81
there are 2 parts to this: 1. harness 2. virtual workspace when you put a good harness together w/ a good workspace experience, it rules. my personal favorites right now are amp/codex/omp. codex's computer use can't be beat if you're building/using your own computer rn.
ai agent/harness tiers
1
1
105
i bet muse will be profitable within the year. once ads get wired up into it it will COOK. the data meta gets (to drive the ad platform) from everyone connecting their lives to it may be positive without even wiring ads into muse directly
Free access to Instinct & Muse feels like $3 Uber Pool rides all over again. Enjoy this stage of the subsidy cycle, we'll be nostalgic for it later 🫡
1
2
151
ive had the same experience with grok 4.7. it's like a mule (workhorse). dependable and doesnt surprise you.
day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless also ignore the reports that compare models with 3d games - that’s not real work. it's made for attention on social media i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i'm ignoring 4.6 because 4.5 has been working better in my experience) key differences with 4.7 - 1. it follows system prompt very, very closely i noticed firstmate showing many new behaviors that i've never seen before, such as asking me to name specific red CI checks that i'm ok with bypassing, and refuse a simple "yolo" instruction i traced it and it's indeed how i instructed it in firstmate's system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up there were a few other similar examples as well. so to me this is a clear behavioral difference 2. it's very "stable" if you've used astra then you know what a "spiky" model is. it can have some genius moments but you occasionally also wonder "how could it be so dumb and doesn't get me". grok 4.7 is the opposite of that throughout the whole day so far, i'll be honest i haven't get a "wow this is absolutely genius" moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly 3. it's a conservative model it doesn't like to take actions without asking, and would explicitly say so this is a bit of a double edged sword, because it means i sometimes have to state the obvious "yes i do want that", but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation 4. it's a bit slower and costs more than 4.5, visibly turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven't quantified exactly where this is coming from yet so overall, i think it's showing some clearly different traits, and i mostly like the changes. i'm going to keep it as my primary firstmate and observe more if you've been using it, what qualitative insights have you gathered from real usage so far?
119
an example of free money: paste this in your bot of choice and see what happens "Help me file my claim in Apple's Siri settlement. Check my eligibility, list what I need, and walk me through every step on the official claim form."
one of my favorite use cases of these ai agents is to "find" money. basically just do the things and follow the processes these companies set up but make too annoying for normal people to follow.
1
1
488
one of my favorite use cases of these ai agents is to "find" money. basically just do the things and follow the processes these companies set up but make too annoying for normal people to follow.
Flight delayed 7 hours - asked Muse to file for compensation. 5 mins later i had $250 credit in my delta account. it even found and rebooked me a new flight. it just figured everything out. even responded to the support email itself. shit feels like magic.
1
1
8
974
happy 4.7 day to those who celebrate
Grok 4.7 is out It brings improvements over 4.6 and is especially good for: >Software engineering >Long-running knowledge work >Electrical engineering >Legal tasks Pricing is the same as Grok 4.6
133
Imagine seeing this graph and thinking it shows close to the full picture. Most closed model tokens (probably >99%) are served direct from the labs or from AWS/GCP/MSFT.
1
168
i hope when grok 4.7 comes out it pushes people to be ambitious instead of telling them to scope down their ambition
1
2
242
One should ask themselves why a16z would compare a SM6 to a … covenant anthem (worst example on this graphic). Surely they are not stupid enough to believe they’re interchangeable. So do they think their investors are idiots? What is it?
One missile, or seven. Same price. a16z's @DefenseInDepth_ on the arsenal the US could have for the price of the one it's buying: a16z.news/p/magazine-depth-a…
174
Had my computer analyze ranked wins in the last 2 years vs $$ spent. interesting to see which teams are outperforming/underperforming
College football’s most carefully guarded secret is the amount of money flowing to the players on each roster. So we reached out to 70 sources across the sport. Estimated budgets for all 68 teams: nyti.ms/4xwoPqe
1
3
198
John Shahawy retweeted
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1,922
2,865
1,539
28,426
7,649,528
a great way to think about process automation: your process is a "loop." once your loop is running at a high enough QA rate. you can plug the entire loop into another (bigger) loop as a single step in the process.
106
muse/facebook is going to introduce normal people to ai 👀
1
98
someone should have told the ai doomers a psyop of the scale they're attempting to pull off requires at least a small amount of charisma
Holy crap POTUS just phoned in @JensenHuang live on stage at All In Summit We will not lose the ai race! And whatever Dario said this weekend won’t stop our progress This made my morning! $NVDA
2
219
i'm happy to pay ~100% more for goods mfg in the USA vs abroad. maybe more. looking at equipment and knowing it was made here hits different.
1
1
128
this is the right way to deal with bottlenecks that users see. thank you @OpenAI
This would suck, but we will prioritize great service for customers until we can get back on top of things.
1
215
Coding on the iPhone duo is going to be so sick
121
Can’t believe OpenAI is cancelling all the $200 plans so they can solve the millennium math problems
1
137