Can Copilot Really Build a Model When You Test It Against the FMI Case Library
Microsoft has made an interesting bet about how to earn credibility with finance professionals: instead of grading Copilot in Excel on generic text tasks, it [partnered with the Financial Modeling Institute](https://cfotech.asia/story/microsoft-taps-finance-body-to-guide-copilot-in-excel), the body that credentials the industry's most demanding modelers, and made FMI's library of real-world modeling cases part of how it evaluates the assistant. The framework judges models on rigour, transparency, flexibility and decision usefulness, which are exactly the axes a reviewer cares about. So the fair question is: held to that bar, can Copilot actually build a model?
The most honest answer available comes from independent testing, and it is "part of one." On the [Excel Modeling Benchmark](https://www.vals.ai/benchmarks/emb), which asks agents to build complete investment-banking and private-equity models from a prompt and source files, the leading model reaches around 69 percent under partial-credit grading, with others clustered just below. The benchmark's own conclusion is the important part: even top scores represent substantial but incomplete models, not client-ready deliverables. Difficulty is consistent across the board, with the highest scores on data-room summaries and the lowest on LBO and DCF work, the most deeply interdependent models, where one early error cascades through everything downstream.
Dig into where the marks are lost and a clear pattern emerges. The leading model passes roughly 87 percent of formula checks and 74 percent of presentation checks but only about 61 percent of numerical checks. In plain terms, it builds a structurally plausible, decently formatted model that quietly gets numbers wrong. A separate [hands-on test by Wall Street Prep](https://www.wallstreetprep.com/knowledge/ranking-the-best-ai-tools-for-financial-modeling-2026/) reached the same practical verdict: the tools are genuinely useful for taking a model from zero to perhaps sixty percent complete, and genuinely dangerous if you trust them to finish the job without extensive review. Even the best still underperformed a junior analyst.
None of that makes Copilot useless for modeling. It makes it a fast, tireless first-drafter that you must never mistake for a reviewer. The workflow that works is to let it frame the structure and populate the mechanical parts, then have a human own every number, especially in the interdependent schedules where errors hide.
At Cell Fusion Solutions Inc we set up exactly that division of labour, using Copilot to accelerate the build while keeping human sign-off on the figures that actually drive a decision. The model it hands you is a strong start. It is not the finish.