Every AI tool we've analyzed since July 2026, distilled into actionable insights for your business. No fluff. Just what works, what doesn't, and what to do about it.
Last updated: October 11, 2026. 27 videos analyzed. 22 from Julian Goldie SEO, 1 from Alex Finn, 4 pending.
What: 8B parameter model that outputs probabilities, not text. Built for routing decisions.
Key feature: Pointer head architecture — computes mathematical relationship between context and options. "Billing 96.41%, Technical 0.10%, Other 3.49%."
How we use it: Agent routing. Instead of prompting an LLM to decide which agent handles a task, Drex gives structured probabilities. Faster, cheaper, more accurate.
Cost: Free (open source, Apache 2.0). API at drex.nace.ai or self-hosted.
Open Source Decision Model Agent Routing
What: One model replaces 9 safety classifiers. Handles routing, PII detection, hallucination detection, and safety checks in one call.
Key feature: Span-level decisions — points to exact words that are problematic, not just "there's an issue."
How we use it: Quality gate for AI Suite agent outputs. Catches PII, unsupported claims, and safety issues before content ships.
Cost: Free (open source, Apache 2.0). 756M params (CPU-friendly) to 9B (GPU).
Open Source Safety PII Detection Hallucination
What: Model trained to investigate AI agent runs and find hidden failures. 0.835 F1 vs GPT-6's 0.816.
Key feature: Finds errors buried in step 47 of 200-step processes. Reads traces like a forensic investigator.
How we use it: Debugging failed cron runs. When a cron fails silently, Flow-1 analyzes the trace and finds the root cause.
Cost: Via Laminar (laminar.sh). Pay-per-use.
Debugging Agent Traces Forensic
What: Predicts the next moment, not the next word. Learned physics by watching video.
Key feature: Self-recovery — robot arm missed a grab, realigned itself, and tried again without being trained on that failure.
How we use it: Not directly applicable to text work, but the self-recovery concept informs our agent error handling.
World Model Robotics Self-Recovery
What: Build playable games by describing them. Play-tweak-play loop with real-time tuning.
Key feature: Tweak panel lets you adjust speed, gravity, damage in real time while playing.
How we use it: The play-tweak-play concept applies to content creation: publish → measure → adjust → republish.
Game Dev Zero Code Play-Tweak-Play
1.05 trillion parameters, 49B active per token. 1M token context. Open weights dropping later this month. Built for real business work — coding, legal, financial analysis, cybersecurity.
Action: Use for complex agent tasks that need long context. The 49B active params make it efficient despite the 1T total.
501B parameters, 23B active. SWEBench 80.9 (ahead of Inkling 77.6 and Neatron 70.7). 3-4x less inference compute than competitors.
Action: Use for coding agents. The efficiency claim means lower costs for the same quality.
1M token context, 128K output. Fastest Claude model. Built for high-volume work — titles, meta descriptions, page checks, keyword groups.
Action: Use for batch SEO tasks. The speed + context window makes it ideal for processing hundreds of pages.
600B parameters, 27B active. 1M token context. Vision support. Free for one week across OpenCode, Kilo Code.
Action: Use as a free sub-agent model. Fast, capable, and free — perfect for background tasks.
Free models, voice assistant, 24/7 news scanner, content ideas, lead gen, video, images — all under one roof. Swap the brain (model) whenever you want.
Action: We already run this. The key insight: one system, multiple agents, each with a job and schedule.
10 AI agents, one screen, zero humans. Jarvis-style. Five layers: founder → hire team → give goal → walk away.
Action: The org chart concept — Ada as CEO, agents reporting to her — is the future of agent management.
AI that builds and fixes its own agents. Reads real conversations, test results, and current setup. Then makes changes itself.
Action: The investigate → change → validate → approve loop is the pattern for self-improving agents.
Forensic investigator for AI agent traces. Finds errors in step 47 of 200-step processes. 0.835 F1 score.
Action: When a cron fails silently, send the trace to Flow-1. It finds the root cause faster than manual debugging.
Search past agent conversations across Claude Code, OpenCode, Grok — all in one place. Indexed for speed.
Action: When you think you solved a problem before but can't remember where, search your agent history.
One model replaces 9 safety classifiers. Span-level decisions show exact words that are problematic. Supports choice, noul, score, span, and set question types.
Action: Use as a quality gate before any AI output ships. The span-level flagging catches issues that simple pass/fail checks miss.
Pulls Google Search Console data. AI visibility report. Keyword opportunity finder. Replaces multiple SEO tools with one system.
Action: Use for keyword research and AI visibility tracking. The GSC integration means no extra API costs.
Google Keep voice input, Skills replacing Gems (Oct 13), Google Vids editing fixes. Skills can be stacked and saved.
Action: Set up Google Skills for your business. They replace Gems and can be reused across sessions.
32GB RAM runs local AI models via Hermes Desktop. Free, no technical skills needed. Can replace cloud subscriptions for some tasks.
Action: If you have a Mac Mini, try local models for tasks that don't need the latest frontier model. Saves on API costs.
Open-source, agent-driven video editing. HTML/CSS/JS-based. Agents write the video, you edit on the same timeline.
Action: Use for AI-generated video content. The agent writes the recipe, the computer cooks the video.
Predicts the next moment, not the next word. Learned physics by watching video. Self-recovery from mistakes.
Action: Not directly applicable yet, but the self-recovery concept is the future of agent error handling.
Zero-code game building. Play-tweak-play loop. Real-time tuning panel.
Action: The play-tweak-play concept applies to content: publish → measure → adjust → republish.
The tools we've analyzed since July 2026 point to one clear trend: AI is moving from "ask a chatbot" to "run a system."
The businesses that win in 2026 and beyond are the ones that:
We've integrated the top 5 tools into AI Suite. The rest are on the roadmap.
🤖 Want this running for your business? Done-for-you Agent OS from £99/month → aisuitehq.org/offer
Drex 1.1 is an open-source decision model that outputs probabilities instead of text. It's built for routing decisions — which team handles a ticket, how urgent is it, does the customer want a refund. It's faster and cheaper than using a language model for the same task.
Vela 2.0 replaces 9 separate safety classifiers with one model. It handles routing, PII detection, hallucination detection, and safety checks in a single call. The key innovation is span-level decisions — it shows you the exact words that are problematic, not just a yes/no flag.
Flow-1 is a model trained to investigate AI agent traces. When an agent run fails silently — especially in long, multi-step processes — Flow-1 can analyze the trace and find the root cause. It's like a forensic investigator for AI systems.
Drex 1.1 and Vela 2.0 are open source and free. You can self-host them or use their APIs. Flow-1 is available through Laminar (laminar.sh). We've built integration scripts for all three — contact us if you want them set up for your business.
A language model generates text. A decision model outputs probabilities. For routing decisions — "which team handles this?" — a decision model is faster, cheaper, and more accurate because it computes the mathematical relationship between context and options rather than generating a paragraph you have to interpret.
Last updated: October 11, 2026 · 27 videos analyzed · ← Back to AI Suite