Blog
How to Evaluate the Investment Research Capabilities of Large Language Models?
When evaluating financial large language models (LLMs), first confirm whether you are testing a bare model, a tool-augmented application, or a complete investment research system. Then examine the question bank: are the facts correct, can the numbers be recalculated, do the sources support the conclusions, and has the model used information beyond its knowledge cutoff date?
Financial LLM evaluation means testing a model or system's facts, calculations, sources, reasoning, and risk boundaries using a verifiable question bank under fixed scenarios, versions, tools, and time points. Model versions, tool permissions, prompts, and raw outputs must also be documented; otherwise, results cannot be verified. Such evaluations can compare research reliability under specific conditions but cannot prove investment returns. This article provides a seven-step process and a six-question test that individual investors can complete in 30–60 minutes.
Key Takeaways
Bare models, tool-augmented applications, and complete investment research systems are three types of evaluation targets, and their results cannot be directly compared across categories.
CFinBench focuses on Chinese financial knowledge; FinEval extends to industry, safety, agents, multimodal, and rigor; BizFinBench.v2 is closer to real user and financial business tasks.
Scoring should cover facts and calculations, sources and timeliness, reasoning and counter-evidence, risk and compliance; stability, speed, and cost should be reported separately.
Serious financial hallucinations—such as fabricating key figures, announcements, or sources—should trigger a zero score for that question or a veto, and must not be masked by average scores.
Repeated generation tests system output stability; repeatedly scoring the same batch of answers only tests the reproducibility of the scoring process.
Before Evaluating: Are You Testing the Model, the Tools, or the Complete System?
Financial LLM evaluation must first define the target. The same model can perform completely differently once connected to search, market data, announcement databases, calculators, and verification workflows. Without layering first, product engineering capabilities can easily be misattributed as base model capabilities.
Three Layers of Evaluation Targets
Bare model: Only allowed to use materials provided in the question; no internet access, no external tool calls. Suitable for testing financial knowledge, reading comprehension, calculation, and reasoning under closed conditions.
Tool-augmented application: The model can call unified search, market data, announcement, or code tools. Suitable for testing tool selection, parameter filling, source verification, time judgment, and multi-step task completion.
Complete investment research system: Beyond the model and tools, it also includes an agent harness, data sources, memory, workflows, retries, permissions, and human review. It answers "can the system reliably deliver research," not "which base model is smarter."
If a team is procuring an AI investment research product, looking only at bare model exam scores is insufficient; if a team is selecting a base model, the proprietary data and workflow advantages of a specific product should not be attributed to the model itself. For more on continuous research workflows, see what is an investment research agent.
How to Evaluate Financial LLMs? Follow Seven Steps
To make evaluations reproducible, at least seven types of records must be kept: questions, configurations, raw answers, tool traces, scores, failure cases, and version changes. Financial LLM evaluation can follow these seven steps:
Fix the scenario and decision problem. First state whether the evaluation is for personal research, content production, research analyst assistance, or product integration. Different use cases have different requirements for depth, timeliness, latency, and risk boundaries.
Fix the evaluation target and category. Specify whether you are evaluating a bare model, tool-augmented application, or complete system. Closed-book vs. internet-connected, API version vs. product version should be reported separately.
Fix the question bank, reference facts, and information cutoff time. Each question should include the task objective, input materials, required answer points, prohibited answer points, source snapshots, acceptable tolerances, and scoring anchors. "Latest" must be replaced with a specific date, time zone, and market status. For data boundaries, see the financial data sources and timeliness guide.
Align configurations. Record the full model version, thinking mode, temperature, maximum output, prompt version, tool permissions, call limits, timeout, and retry rules. Conditions that cannot be unified should be evaluated in separate groups.
Generate repeatedly and preserve the full process. Run the same question multiple times, saving complete answers, citation links, tool traces, token usage, time elapsed, error types, and failed retries. Keeping only the best run overestimates real-world usability.
Score in layers and enforce red-line arbitration. For objective calculations, prefer programmatic or standard-answer checking; for open research questions, use anonymous model judges and human review. Serious financial hallucinations should be handled separately.
Disclose evidence and limitations. At minimum, publish the methodology card, sample questions, raw output samples, scoring anchors, failure cases, exclusion rules, version logs, and conflicts of interest.
The most dangerous state of a financial answer is not that it is short, but that it is complete, professional, and readable, yet built on incorrect facts or incorrect time points. Therefore, quality scores should not only reward coverage, and efficiency metrics should not be blended with financial reliability into a single composite score.
Metric | Main Checks | Recommended Judgment Method | Common Misjudgments |
Facts & Calculations | Numbers, formulas, units, YoY/QoQ, net profit definitions, fiscal year definitions | Standard answers, verifiable calculations, tolerance rules | Ignoring final numerical errors because the derivation is long |
Sources & Timeliness | Whether sources exist, are primary, publication date, data cutoff date, whether citations support assertions | nashnova's research evidence chain method: assertion → citation → source → original text excerpt | Treating any link as evidence; old data written as latest |
Reasoning & Counter-evidence | Causal chains, alternative explanations, key assumptions, falsification conditions, and information gaps | Clear scoring anchors, anonymous dual judges, human spot-checks | Only rewarding smooth conclusions without checking opposing evidence |
Risk & Compliance | Uncertainty, investor suitability (whether the product and risk match the user), leverage and liquidity risks, whether it oversteps by giving directives | Red-line checklist and financial human review | Earning risk scores just by piling on disclaimers |
Stability | Variation in facts, conclusions, and scores across multiple runs of the same question | Report mean, dispersion, and failure rate | Confusing re-scoring the same answer with repeated generation |
Speed & Cost | Latency, token usage, tool call count, API cost, and timeout rate | Separate efficiency table | Low cost masking factual errors |
Specific weights should follow the use case. nashnova's public evaluation does not force-combine three systems into a single composite score: financial reliability focuses on facts, timing, policy nature, calculations, evidence, and red lines; research quality uses five equally weighted dimensions—data coverage, data accuracy, reasoning quality, conclusion quality, and answer quality; practical usage experience separately examines facts and calculations, stability, length, wait time, and batch task suitability.
Consider a simplified example: Report A has broader coverage but writes a draft-for-comment regulation as an enacted policy; Report B is shorter but clearly labels the policy status, sources, and falsification conditions. When the rubric emphasizes report completeness, A may lead; when it emphasizes financial reliability, B is more reasonable. Ranking changes do not necessarily come from model changes—they may come from changes in scoring objectives.
What 10 Items Should Individual Investors Check When Reading Financial LLM Evaluation Reports?
For individual investors, when you see a "financial LLM ranking," don't look at the top spot first. Check the following ten items; when core items such as evaluation target, information cutoff time, raw outputs, or scoring rules are missing, you should significantly lower your trust in the conclusions.
Evaluation target: Bare model, tool-augmented application, or complete system?
Scenario and audience: Serving individual investors, research analysts, content teams, or product integration?
Question bank structure: Are question volume, task types, difficulty levels, adversarial questions, and public question ratios disclosed?
Time point and sources: Are the information cutoff time, primary sources, source snapshots, and retrieval dates verifiable?
Configuration fairness: Are model versions, prompts, thinking modes, tool permissions, timeouts, and retries aligned?
Raw outputs: Are both successful and failed samples disclosed, rather than only curated screenshots?
Scoring rubric: Are weights, anchors, judge identities, blinding, and human review rules transparent?
Red-line handling: Are serious hallucinations, future information leakage, time-point errors, and compliance violations listed separately?
Stability and efficiency: Were outputs truly generated repeatedly, with time, cost, and failure rates reported separately?
Limitations and conflicts of interest: Who organized the evaluation, what was excluded, what can it prove, and what can it not prove?
A credible report does not necessarily need to disclose all proprietary questions, but it should provide sufficient methods and evidence for verification. At minimum, one should be able to trace from published conclusions back to the corresponding questions, answers, scoring rules, and sources.
How to Interpret nashnova Investment Research Framework Evaluation Results?
The nashnova investment research framework places six models into the same set of financial data, investment research capabilities, and agent workflows, observing financial reliability, research quality and depth, and practical usage experience separately. The three dimensions address different questions and should be read independently.
Evaluation Dimension | Primary Focus | Suitable for Judging | Should Not Be Extrapolated As |
Financial Reliability | Facts, time points, policy nature, calculations, sources, and critical errors | Whether the model requires stricter review in high-risk financial tasks | Investment returns or absolute safety |
Research Quality & Depth | Data coverage, data accuracy, reasoning, conclusions, and answer quality | The model's content ceiling when completing complex research tasks | Stable performance across all question types |
Practical Usage Experience | Facts and calculations, output stability, wait time, and batch task suitability | Whether the model is suitable for daily research workflows | Absolute capability ranking of the API bare model |
This evaluation compares the actual performance of models after they are integrated into the nashnova investment research environment. It cannot replace investors' own verification, nor can it prove future returns.
Conclusion: Build the Evidence Chain First, Then Discuss Rankings
Evaluating financial LLMs is not about running a set of questions and checking the total score. First fix the target, scenario, version, and tool permissions, then check facts, calculations, sources, time points, and risk boundaries; for open-ended questions, use blinded judges and human review, and handle serious hallucinations separately according to pre-published rules.
Individual investors can use six questions to first eliminate obviously unreliable tools. Product and research teams need to retain question bank versions, raw outputs, tool traces, and scoring records so that every conclusion can be verified. Rankings can be referenced, but only after the evidence chain is established.
