AI in Finance: Claude Fable 5 Tops at 49% on Investment Tasks

5d ago·0:00 listen·Source: Tech Times

Summary

New evaluation results show that top AI models are not yet ready to handle complex financial research. The FrontierFinance benchmark, designed by Samaya AI, tested AI agent performance across the full investment workflow. Anthropic's Claude Fable 5, the top-performing frontier model, completed just 49.2% of the benchmark's tasks. OpenAI's GPT-5.6 Sol scored 46.8%, and Claude Opus 4.8 achieved 45%. All tested models, including Google's Gemini 3.1 Pro, fell below the halfway mark. Samaya's own proprietary system achieved 50.8% accuracy in a "low effort" configuration and 56% in a "high effort" mode. The benchmark covers tasks like synthesizing sector-wide information, screening for opportunities, and tracking covenants through earnings commentary. This matters because these are the tasks that drive real capital allocation decisions in finance.

Read the full article on Tech Times

This is an AI-generated audio summary. Always check the original source for complete reporting.

Share
Keep Listening