Evaluations and notes on AI agents across real-world use cases.
Research
-
Dont vibe code your software clones
Why high-fidelity software replicas demand more than surface-level generation.
-
Evaluating Sovereign AI
Testing Sarvam models on multilingual ecommerce support tasks with tools, policy constraints, and backend state.
-
Tech Support Environment
The Tham Luang Cave
-
SalesforceBench
Can agents actually work inside a simulated Salesforce org?
-
Editing is Hard
Can LLMs edit PPTX reliably?