Download nashnova App

Blog

DeepSeek V4 Pro: Can It Become the Go-To Model for Daily Investment Research?

On August 13, DeepSeek V4 Pro was officially released.

Plenty of model benchmarks have already appeared on the market, but for individual investors, leaderboards focused on coding ability, general knowledge Q&A, or a single composite score only provide partial reference. What truly matters for investment research is whether a model can read financial statements, handle valuation calculations, explain market logic, and avoid critical errors when dealing with facts, timelines, and policy classifications.

With this question in mind, nashnova connected DeepSeek V4 Pro, Claude Opus 5, Qwen3.8-Max, DeepSeek V4 Flash, Kimi K3, and GLM-5.2 to the same investment research product environment, using identical financial data, research Skills, and Agent workflows. Three separate evaluation suites were conducted: financial reliability, research quality & depth, and real-world usability.

Open-ended research questions were scored with the participation of GPT 5.6 flagship model sol, so that model was excluded from this head-to-head comparison.

We aimed to answer one specific question:

Can DeepSeek V4 Pro serve as the primary model for daily investment research?

The conclusion: yes, but it is not the best choice for every task.

DeepSeek V4 Pro is well-suited for valuation, standardized research, and batch tasks; Claude Opus 5 performs better in complex company research and deep reasoning; Qwen3.8-Max is more suitable for policy, macro, and event analysis; DeepSeek V4 Flash is ideal for quickly screening large volumes of news, earnings reports, and industry materials.

Kimi K3 has advantages in organizing long documents and industry information, but timelines and policy facts require additional verification. GLM-5.2 can produce complete analytical structures, but in this round of testing it had relatively more critical factual and calculation errors, so it is not recommended for important investment research tasks or financial content publication at this time.

Why Three Separate Evaluation Suites Are Needed

An investment research answer can be very deep yet contain errors in key figures; it can also rarely make mistakes but keep its reasoning at a superficial level. Some models are good at quickly processing large volumes of tasks but are not suitable for delivering the final output of high-stakes research.

If these differences are compressed into a single composite score, the result can easily mislead users. Therefore, nashnova split the evaluation into three questions:

Evaluation Suite

What It Primarily Checks

What Question It Answers

Financial Reliability

Facts, timelines, policy classifications, calculations, evidence, and critical errors

Can the answer enter a formal research or publication workflow?

Research Quality & Depth

Data coverage, data accuracy, reasoning, conclusions, and presentation quality

Can the model complete complex research and produce a workable draft?

Real-World Usability

Factual & calculation quality, output length, wait time, and batch-task suitability

Is the model suitable for daily, repetitive research work?

The three suites use different criteria, and their rankings do not fully align. This is not a contradiction—it shows that model selection must start from the task at hand.

Daily Use: V4 Pro Is Well-Suited for Standardized and Batch Tasks

In the real-world usability evaluation, DeepSeek V4 Pro and Claude Opus 5 had similar consensus scores.

V4 Pro's strengths are mainly in valuation, quantitative estimation, and standardized research. It typically lists formulas, units, and calculation steps clearly, making subsequent verification easier. For batch earnings Q&A, initial company screening, or valuation tasks under fixed templates, this delivery style is more practical than occasionally producing a beautifully written answer that is hard to replicate consistently.

It should be noted that the real-world usability scores are consensus scores derived from two independent scoring rounds. Two scoring rounds do not mean the model regenerated two batches of answers under a frozen configuration, so this result alone cannot prove the model's generation-to-generation stability. A more accurate statement is: within this evaluation sample and scoring framework, V4 Pro demonstrated strong daily task suitability and standardized delivery capability.

Claude Opus 5's advantage emerges in questions requiring deep analysis. It more frequently proactively identifies key assumptions, counter-evidence, falsification conditions, and follow-up tracking indicators, making it suitable for in-depth company research, industry breakdowns, and complex scenario analysis. The trade-off is that its answers are usually longer and take more time.

Qwen3.8-Max has relatively high information density and performs well in policy, macro, and event analysis. Users should still double-check amounts, units, and ratios to prevent calculation details from affecting conclusions.

DeepSeek V4 Flash has a clearer positioning. It is suited for processing large volumes of news, earnings reports, and industry materials first, then passing the truly important questions to a stronger model for further research.

Kimi K3 is better suited for long documents and industry information organization, but chronological order, policy status, and key facts need verification. GLM-5.2 typically produces complete analytical structures, but had relatively more underlying factual and calculation errors in this round of testing.

Financial Reliability: Smooth Writing Doesn't Mean No Errors in Critical Areas

Financial research cannot be judged solely on whether an answer is complete and fluent. A single critical factual error can invalidate the entire downstream analysis chain.

Several typical issues appeared in this round of testing: using information published later to explain earlier market movements, misclassifying regulatory or licensing fees as tariffs, and writing 800 million yuan as 8.1 billion yuan.

nashnova defines errors severe enough to undermine core conclusions as "red-line errors." These errors cannot be offset by high scores in other areas.

Across the entire test batch, 5 red-line errors were confirmed: Qwen3.8-Max and Claude Opus 5 triggered none; DeepSeek V4 Flash, DeepSeek V4 Pro, GLM-5.2, and Kimi K3 each had one or more.

This also explains V4 Pro's seemingly contradictory results. It performed well in daily use and standardized tasks, but did not rank at the top in the financial reliability evaluation, because one of its answers misclassified regulatory or licensing fees as tariffs.

This is not a wording issue. Once the nature of a policy is misidentified, the downstream reasoning about industry costs, corporate profitability, and future policy direction can all deviate from reality.

Qwen3.8-Max led the financial reliability evaluation, primarily because it was more stable when handling policy classifications, event timelines, and formal content. Claude Opus 5 was also in the top tier while maintaining high research depth. Regardless of model rankings, however, key figures, sources, and calculations still require manual verification.

Research Depth: Claude Comes Closest to a Workable Research Draft

Avoiding critical errors only addresses the baseline of research. Completing company research, industry analysis, or scenario modeling also requires the model to develop assumptions, counter-arguments, and follow-up observation items.

In this area, Claude Opus 5's advantage was most apparent.

It typically explains what conditions a conclusion depends on, what evidence could overturn the current judgment, and what indicators should be tracked going forward. Compared to simply providing an answer, this structure is closer to a research draft that can be further verified and iterated upon.

Qwen3.8-Max provides relatively high information density and is suitable for drafting policy, macro, and event research, though key calculations still need separate verification.

DeepSeek V4 Pro has a more structured format, suitable for completing valuation, company Q&A, and standardized research under fixed templates. When it comes to multi-layered reasoning, counter-arguments, and complex scenario analysis, Claude achieves a higher degree of completeness.

Kimi K3 has certain advantages in organizing industry materials. DeepSeek V4 Flash is suitable for preliminary screening but should not be directly relied upon for complex research final drafts.

Putting all three sets of results together, V4 Pro's role becomes clear: it can handle most standardized daily research work, but for tasks requiring depth, policy analysis, or high-stakes decisions, users should switch models or apply stricter cross-verification.

How Individual Investors Should Choose Models by Task

Users do not need to subscribe to all models because of this evaluation. A more practical approach is to establish a primary model and a verification model based on your most frequent tasks.

Primary Task

First Choice

Verification / Supplement

Usage Recommendation

In-depth company or industry research

Claude Opus 5

Qwen3.8-Max

Claude produces the research draft; Qwen checks policy, facts, and timelines

Policy, macro, and event analysis

Qwen3.8-Max

Claude Opus 5

Cross-reference against original policy documents; double-check key figures

Valuation and quantitative estimation

DeepSeek V4 Pro

Claude or spreadsheet tools

Retain formulas, units, and denominators; use Excel or similar tools to recalculate

Daily batch research

DeepSeek V4 Pro

DeepSeek V4 Flash

Flash screens materials first; V4 Pro handles key questions

Long documents and industry materials

Kimi K3

Claude Opus 5

Kimi organizes the structure; Claude digs into core issues

Formal publication or high-stakes tasks

Claude or Qwen

Cross-verify with another model

Final confirmation by a human against announcements, policy documents, and raw data

For important tasks, a simple dual-model workflow can be adopted: one model handles the analysis, another model only checks facts, timelines, units, denominators, policy classifications, and calculations, with a human making the final confirmation against original announcements or policy documents.

After the analysis is complete, you can append the following prompt separately:

Please independently check facts, timelines, units, denominators, policy classifications, and whether citations support the conclusions. Only list errors that could overturn the conclusions—do not continue polishing the main text.

This prompt cannot replace manual verification, but it can help users more quickly identify issues that could change conclusions.

Why nashnova Does Not Provide a Unified Overall Ranking

This evaluation compares the actual performance of six models after being connected to nashnova's investment research environment, not their capabilities on official chat pages or as bare API models.

Within the same product environment, models have access to the same financial data, research Skills, and Agent workflows. The purpose is to reduce differences in data and workflows so that results more closely reflect users' experience in real investment research tasks.

However, this also means the evaluation results reflect the combined performance of the model and the product environment, and cannot be directly extrapolated as absolute rankings of base models across all scenarios.

nashnova's product philosophy is also not about having users bet on a single model that is forever ranked first. Different research tasks have different requirements for reliability, depth, speed, and cost. Users should switch models based on the task and always retain source verification, calculation checks, counter-argument review, and human confirmation.

So, returning to the opening question: Can DeepSeek V4 Pro serve as the primary model for daily investment research?

Yes.

It is well-suited for valuation, standardized research, and batch tasks, and can handle most daily research work. But "primary" does not mean "all-purpose." For complex research, prioritize Claude; for policy and macro questions, prioritize Qwen; for large volumes of materials, let Flash handle the first pass. Important conclusions still require cross-verification and human confirmation.

Appendix: Full Results of All Three Evaluation Suites

1. Financial Reliability Evaluation

This suite focuses on whether answers can enter a formal research or publication workflow. Scoring covers facts, timelines, policy classifications, calculations, evidence, reasoning, and presentation. If a critical error severe enough to overturn conclusions occurs, the relevant question's score is capped.

Rank

Model

Quality Score

Final Score

Red-Line Count

1

Qwen3.8-Max

87.65

87.65

0

2

Claude Opus 5

85.89

85.89

0

3

DeepSeek V4 Flash

85.71

81.15

1

4

DeepSeek V4 Pro

84.83

80.57

1

5

GLM-5.2

83.45

78.59

1

6

Kimi K3

82.29

75.29

2

The quality score reflects overall performance before red-line processing; the final score incorporates the cap applied for critical errors. Qwen3.8-Max and Claude Opus 5 triggered no red lines; the other models' quality scores were not low, but critical errors dragged down their final results.

2. Research Quality & Depth Evaluation

This suite uses five equally weighted dimensions: data coverage, data accuracy, reasoning quality, conclusion quality, and answer quality. Each dimension has a maximum of 5 points, with a per-question maximum of 25 points. The official total score is based on 13 blind-evaluated questions.

Rank

Model

Total Score

Score Rate

Per-Question #1

1

Claude Opus 5

323 / 325

99.4%

11

2

Qwen3.8-Max

292 / 325

89.8%

6

3

Kimi K3

282 / 325

86.8%

2

4

DeepSeek V4 Pro

269 / 325

82.8%

0

5

GLM-5.2

261 / 325

80.3%

1

6

DeepSeek V4 Flash

248 / 325

76.3%

0

Claude Opus 5's lead in complex research is quite pronounced. The near-perfect result also reminds us that a single model judge and its writing preferences may influence open-ended question scoring, so this table is better suited for observing research ceilings rather than representing all investment research capabilities on its own.

3. Real-World Usability Evaluation

This suite focuses on the overall user experience of models within the product, including factual & calculation quality, output length, wait time, and batch-task suitability. Scores are consensus scores derived from two independent scoring rounds.

Rank

Model

Consensus Score

Usage Positioning

1

Claude Opus 5

85.9

Higher research ceiling; longer output

2

DeepSeek V4 Pro

84.4

Suited for daily and standardized tasks

3

Qwen3.8-Max

80.7

Strong overall capability; calculations need verification

4

Kimi K3

80.3

Good for industry and long-document tasks

5

DeepSeek V4 Flash

77.0

Speed-first option

6

GLM-5.2

61.0

Higher tail-end error risk

Some policy and cross-asset questions were separated out and not included in this general capability main table. Therefore, it is suitable for assessing daily usability but not for independently judging policy task risk, nor can it prove generation-to-generation stability.

The three tables should be read separately: financial reliability answers "Can the content enter a formal research workflow?"; research quality answers "Can complex questions be explored in depth?"; real-world usability answers "Is it convenient for daily use?"

This article is for model capability research and educational purposes and does not constitute investment advice. It does not provide stock recommendations, buy/sell instructions, return guarantees, or automated trading services. The release date and official capability information for DeepSeek V4 Pro are based on publicly available materials as of August 13, 2026.