GrokBot and Jev
Joseph Fluckiger
Hi, I'm Joseph.
- I build AI fraud agents at eBay, on LangGraph
- I've run OpenClaw since it launched
- Hermes lives on my M1 Mac
- And now: Grok Bot
Personal AI agent swarms: productivity nirvana or overhyped busywork?
Everyone is converging on the same shape
Always on
Works while your laptop is closed
Own computer
Browser, terminal, files
Chat interface
Desktop, mobile, voice
Skills
Reusable instructions and tools
Routines
Scheduled or triggered work
Same pattern. Different hosting, model choices, and trust.
| Agent | Price | Runs on | Pick the model? | Your data | Chat |
|---|---|---|---|---|---|
| Grok Bot | Bundled, $30+ | Its cloud VM | No | Vendor holds it | Own app, no iMessage |
| Hermes | Free + tokens | Your machine | Any | Yours | 20+ incl. iMessage* |
| OpenClaw | Free + tokens | Your machine | Any | Yours | iMessage, WhatsApp, more |
| Also in the race | |||||
| OpenAI Dots | In ChatGPT Pro | OpenAI cloud | OpenAI | Vendor holds it | App, Slack, Teams |
| Meta Muse | Free; paid ~$20 / ~$100 | Meta VM | Meta | Vendor holds it | App, WhatsApp |
| Claude Cowork | In Pro $20+ | Anthropic cloud | Claude | Vendor holds it | App, Slack |
| Instinct | Invite-only, price not public | Vendor cloud + your devices | Vendor's own | Vendor holds it | iMessage, WhatsApp, calls |
* via a bridge (BlueBubbles). “Free” = free software; you still pay for models and hosting. Prices change weekly. Detailed pricing is in the extra material.
How I use Grok Bot
- #1 Calendar + email Pull an event from calendar/email, enrich it, add only if no conflict
- This day in history Ported from my original OpenClaw agent
- Obsidian vault Personal CRM, projects, draft blog posts
- Strava Auto thumbs-up friends’ workouts; photo → bot → attach to a run
- Coaching Bot Fitness / coaching Grok Bot
Which setup fits you?
| GrokBot | Hermes | OpenClaw | |
|---|---|---|---|
| Hosting | Managed cloud computer | Your computer or VM | Your computer or VM |
| Model choice | Platform chooses | You choose the provider | You choose the provider |
| Messaging | Desktop and mobile app | Chat apps, iMessage via bridge | iMessage, WhatsApp, more |
| Setup | Create a bot, connect apps | Install, configure, maintain | Install, configure, maintain |
| My trade-off | Convenience / vendor dependence | Model freedom / upkeep | Messaging reach / upkeep |
Self-hosting still shares prompts with any cloud model you use. Full comparison and pricing: Extra material.
JEV
Jev is not an LLM
“It doesn’t actually write any text. It actually answers questions by picking from a list and then giving a probability for each one of those answers.”
- Input: State + multiple Questions at once → Output: probabilities over predefined options
- Those probs are real — tied to the categorization, not after-the-fact “I’m 90% sure” text
- Faster & cheaper for judgment calls; answers return all at once. Not good at math/counting.
- LLM “confidence” is generated text — doesn’t match real token probs. RLHF makes models sound sure when wrong.
- Jev trains with RLCD — Reinforcement Learning for Calibrated Decisions.
Named for Jevons Paradox
- William Stanley Jevons (1865): when steam engines got more efficient, Britain used more coal — not less.
- Same idea for System 1 AI: make a judgment call super fast and super cheap → usage explodes.
“When a judgment call is super fast and super cheap like with Jev, well it can go into places where an LLM would be too slow or too expensive today, like in every row of a database or on every line of a log file.”
Also: System 1 vs System 2
System 1 · fast
Quick judgment calls. “What’s 2×2?” — you just know.
Jev = TypeSafe’s System 1 model.
System 2 · slow
Deliberate chain-of-thought. “What’s 17×24?” — work it out.
LLMs / reasoning models ≈ System 2.
TypeSafe borrows Kahneman’s terminology from Thinking, Fast and Slow. This is the classification, not the product name.
Calibrated confidence is a control surface
Automate / act
Defer to a human
Ignore / skip
Raise the bar when a mistake is expensive.
- Guardrail: check jailbreaks before the LLM — or check LLM output before the user
- Agents: route, defer-to-human, or gate actions with real thresholds
LangGraph
LangGraph enables Jev in agentic flow
Code
Exact checks, arithmetic,
caches, aggregation
Jev
Intent, relevance,
bounded judgments
LLM
Drafts, explanations,
open-ended reasoning
Tools
Search, data access,
approved actions
Questions?
Thank you, AIMUG. 🙏
Sources and demo material
- GrokBot: Overview and approvals
- Jev: System One, primitives, confidence, models, API
- RAG: Passage decisions and reranking benchmark
- Speed: 13-question batching benchmark
- Limits: Jev’s documented failure modes
- Self-hosted agents: Hermes and OpenClaw
- Demo kit: Python runner and measurement sheet
- Backup: Extra material
- Talk / video: youtu.be/YGgNBcIgI4s — IBM / Martin Keen lightboard on Jev (TypeSafe)
- Name: William Stanley Jevons (1865) — Jevons Paradox (efficiency → more coal use). Product named for this, not for Kahneman.
- System 1 / 2 terminology: Daniel Kahneman, Thinking, Fast and Slow (2011) — TypeSafe’s classification framing only
- Docs (reported): jev-ai.org/docs — decision API, question types, calibrated answers
Checked Oct 5, 2026. Vendor benchmarks, proposed designs, and live measurements are identified separately.