Andrew Ng said "Our eval tools aren’t ready for LLMs."
@LangWatchAI Evaluations Wizard solves this by simulating real-world interactions and running 30+ evaluators on your LLM app.
Works even if you have no eval dataset!
100% open-source. https://t.co/osImmD7bmw