2026-06-22
LoopCraft Introduces Built-in A/B Testing for Agent Loops
Compare two agent configurations side-by-side with automatic statistical analysis. Test different prompts, models, and loop patterns to find the optimal design.
← Back to NewsWe are excited to launch LoopCraft A/B Testing — a built-in experimentation framework for AI agent loops. Define two agent configurations (different prompts, models, or loop patterns), set your evaluation criteria, and LoopCraft runs both configurations against the same inputs. The dashboard shows side-by-side comparisons with statistical significance indicators — win rate, average latency, token cost per run, and success rate. Bayesian analysis determines when a winner is statistically significant, so you're not fooled by small sample sizes. Export test reports as PDF for stakeholder reviews. You can test any component in isolation: prompts, model selection (GPT-4o vs Claude vs Gemini), loop patterns (ReAct vs Plan-Execute), tool configurations, or even entire agent pipelines. A/B Testing is available on Team and Enterprise plans.