What a multi-agent research app is
Ask a single chatbot whether a stock is worth buying and you get one voice, one pass and one set of blind spots. If the model leans towards a view early in its answer, the rest of the answer tends to support it. Nobody in the conversation is paid to disagree.
A multi-agent research app splits the job up. Several instances of a large language model (LLM) are each given a narrow role, the same underlying data and a different instruction. Some research. Some argue. One decides. The structure borrows from how an investment team divides work: an analyst who reads the accounts, another who reads the chart, someone who tracks the news, and a portfolio manager who has to weigh them.
The agents are not separate intelligences. They are usually the same model, or a few models, prompted differently. What changes is the process: each view is produced separately and then has to survive being challenged.
The pipeline: research, debate, synthesis
Designs vary, but most follow three stages.
- Data retrieval. You enter a ticker. The app pulls in price history, financial statements, recent news and sometimes macro data. Everything downstream depends on this step being current and correct.
- Specialist research. Each specialist works from its own slice:
- a fundamentals agent reads revenue and earnings trends, margins, the balance sheet and what the price implies about future growth;
- a technical agent reads the price history: trend, momentum, support and resistance;
- sentiment and news agents read recent headlines and the tone around the name, and flag events the financials do not yet reflect.
- Adversarial debate. A bull agent and a bear agent each take the specialists’ reports and build the strongest case for one side, then attack the other side’s case point by point. This may run for one round or several.
- Synthesis. A portfolio-manager agent reads the reports and the debate and writes a structured verdict: a view, how confident it is, what would change it, the catalysts ahead and sometimes a suggested position size.
The important design choice is the order. Independent research comes first, argument second, so the specialists are not simply echoing each other. And the bear agent’s whole job is to find what is wrong with the bull case, which a single model answering a single question has no reason to do.
Why debate can help, and what the research shows
The appeal of debate is that it builds in a counter-argument. A model asked to defend a conclusion will usually find reasons to; a second model instructed to attack it will usually find some too. The portfolio manager then sees both, which is closer to how a careful human would want to think about a position.
There is academic work behind the idea, and it is more mixed than marketing usually suggests:
- Du, Li, Torralba, Tenenbaum and Mordatch, Improving Factuality and Reasoning in Language Models through Multiagent Debate (arXiv, 2023), had several model instances propose answers and debate them over multiple rounds. They reported better mathematical and strategic reasoning and fewer false answers and hallucinations on the tasks they tested.
- Smit, Duckworth, Grinsztajn, Barrett and Pretorius, Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs (arXiv, 2023, revised 2024), compared debate methods with other prompting strategies and found that debate, in its then-current forms, did not reliably beat simpler approaches such as self-consistency and ensembling. It could do better with careful tuning, but was sensitive to settings and hard to optimise.
- Xiao, Sun, Luo and Wang, TradingAgents: Multi-Agents LLM Financial Trading Framework (arXiv, first posted December 2024), applied the idea to trading, with fundamentals, sentiment and technical analysts, bull and bear researchers, a trader and a risk-management team. The authors reported better cumulative return, Sharpe ratio and maximum drawdown than their baselines. Read the details, though: the backtest ran from 1 January to 29 March 2024 on a small number of large US technology stocks, and the authors say they limited it to three months because each prediction needed 11 LLM calls and more than 20 tool calls. They also note that the highest Sharpe ratio exceeded the range they expected.
The fair reading is that debate is a sensible way to reduce the risk of one model agreeing with itself, and that the evidence it produces better investment decisions is early, short and narrow. Three months of backtest on a handful of stocks is a starting point for research, not proof of an edge.
Worked example: checking what an agent tells you
The fundamentals agent says the stock “trades at 18 times earnings”
You check the quote: the share price is $120. The company’s last reported annual earnings per share were $5.00.
P/E = price ÷ EPS = $120 ÷ $5.00 = 24
That is not 18. Work backwards to see what the agent must have used:
implied EPS = $120 ÷ 18 = $6.67
$6.67 is not the reported figure, so the agent either used a forward estimate without saying so, pulled a stale or wrong number, or made one up. The first is a labelling problem; the other two are errors. Either way, a bull case built on “cheap at 18 times” looks different at 24 times trailing earnings, and the verdict that rested on it needs re-reading.
The cost side deserves the same arithmetic. Using the TradingAgents paper’s figure of 11 LLM calls per prediction, a 20-stock watchlist run once is 20 × 11 = 220 calls. Run it every trading day for a 21-day month and that becomes 220 × 21 = 4,620 calls. Any given app will make a different number of calls, but the point holds: cost scales with tickers × agents × debate rounds × how often you run it, and on a paid API every one of those calls is billed.
The honest limits
- Hallucination does not go away. Every agent is still an LLM. Debate can catch some errors, but two agents can also argue confidently about a number neither checked. A persuasive debate is not the same as a correct one.
- Data can be stale or wrong. The agents only know what the app retrieved. A missed earnings release, a delayed quote or a news feed that skipped a story flows straight into every downstream view.
- Same model, correlated mistakes. If every agent runs on the same model, they share its blind spots. The bear agent may not think of the objection the model never learned to raise.
- No guarantee, and short evidence. As above, the published results are short backtests on a few stocks. Nothing about the method guarantees future returns.
- An LLM is not an adviser. It does not know your finances, goals, tax position or tolerance for loss, and it is not licensed to give personal advice. A “suggested position size” from a model is a generic starting point, not an instruction.
- API calls cost money. With a bring-your-own-key app you pay the model provider directly, at their rates, for every run.
How to use one responsibly
Treat the output as research notes from a fast, tireless, occasionally wrong junior analyst. That means:
- Verify every number that matters against the company’s filings or a reliable data source before it influences a decision, as in the example above.
- Read the bear case first. It is the part a single-model answer usually leaves out, and the part you are most likely to skip.
- Rebuild the valuation yourself. If the verdict leans on the stock being cheap or expensive, put your own assumptions into a DCF calculator and see whether the conclusion survives reasonable changes to growth and the discount rate.
- Size from risk, not from the verdict. Conviction does not set a stop. Position sizing explained covers turning a stop distance into a share count, and if the position adds to market exposure you already carry, beta-weighted hedging shows how much index hedge offsets it.
- Treat it as research, not a signal. Do not wire an AI verdict to an order button.
HedgeDesk is a multi-agent research app of this kind, for Mac, iPhone, iPad and Android. Fundamentals, technical, sentiment and news analysts research a stock or ETF, bull and bear advocates argue the case, and a portfolio manager agent returns a verdict with conviction, risk factors, catalysts and a suggested position size. You choose the engine: on supported Macs it can run on-device with Apple Intelligence, with no API key, or you connect your own key for a cloud provider and pay that provider directly. Market data comes from Yahoo Finance and macro data from the IMF. The free tier includes Instrument Analysis; HedgeDesk Pro (US$5.99 a month or US$59.99 a year) adds DCF & WACC models, portfolio modelling, the Catalyst Scanner, the Playbook journal and the Growth Engine. Like any tool in this category, its output is for information only and is not financial advice.