Evaluations and notes on AI agents across real-world use cases.

Research

  1. Dont vibe code your software clones

    Why high-fidelity software replicas demand more than surface-level generation.

  2. Evaluating Sovereign AI

    Testing Sarvam models on multilingual ecommerce support tasks with tools, policy constraints, and backend state.

  3. Tech Support Environment

    The Tham Luang Cave

  4. SalesforceBench

    Can agents actually work inside a simulated Salesforce org?

  5. Editing is Hard

    Can LLMs edit PPTX reliably?