Lobsters AI
A Prolog library for interfacing with LLMs
Comments
Model evaluation
Catalog
89 results for Model evaluation
Comments
Comments
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
Comments
Comments
Comments