Capacity planning for consulting firms fails because the pipeline forecast and the staffing decisions live in separate systems, updated on separate cadences, by people who are rarely in the same room. The fix is a system that merges the two, plus the weekly 30-minute operating habit that system makes cheap: one place where the sales pipeline and services delivery sit together, so the questions a human would otherwise agonize over have their inputs side by side. When that runs on clean, merged data, capacity planning compresses from a three-day model exercise into something fast enough to catch a deal that closes on Tuesday.
At most teams under 200 people, the symptom looks like this: the CFO asks whether the firm will make Q3, operations runs the model, the model says 91% utilization, and three weeks later the team is short two senior consultants for a deal that closed unexpectedly. The model was answering the wrong question. At 91% scheduled utilization the answer was never that the team is fine, it was hire now, because there is no slack left to absorb a deal that closes early. The model reported a number; nothing connected that number to the pipeline that would consume the last 9% within weeks. A system that merges pipeline and delivery surfaces the need to hire before the deal closes, not three weeks after.
What capacity planning actually is
Capacity planning for consulting answers one question: do we have enough of the right people, in the right weeks, to deliver the pipeline we are forecasting? If not, the firm is hiring, contracting, or re-forecasting. That is the whole job. Three conditions must be true for the answer to be useful: the pipeline is weighted by stage (not gut-percentaged), the people side is tracked by skill and by week (not headcount by month), and the people who own staffing decisions are in the room when the answer is produced.
When all three conditions are true, capacity planning compresses from a three-day model exercise into a short weekly meeting that produces decisions. When any of them is false, it produces a slide.
Why better forecasting tools and a better operating habit are the same fix
The capacity model and the pipeline forecast live in separate files and never connect. The gap between them, typically three to six weeks from the moment a deal will close to the moment a person is staffed, is part data problem and part decision problem, and the two reinforce each other. Separate files mean the inputs never sit side by side, so the decision never gets made on time. Firms that solve capacity planning change both at once: they merge the pipeline and delivery into one system, and they run the operating cadence that system makes fast.
The common pattern is easy to see. Ask a firm to show its capacity model and someone opens a spreadsheet. Ask for the pipeline forecast and someone opens a different spreadsheet, or a CRM report. Ask where the two connect, and there is silence.
The two never connect because they describe different stages of the same work, and naming those stages makes the gap obvious. The pipeline forecast is the Forecast: deal stage, weighted, for work that has not closed yet. The capacity model is built from the Actual: named-person assignments, which only exist once work is signed. Between them sits the Estimate: your capacity by skill and by week, the supply you could commit if the forecast closed. Capacity planning is the discipline of reconciling the Forecast against that Estimate before the Actual is forced on you. Resource management picks up at the Actual, the named, signed-off assignment, so the two disciplines share one vocabulary and hand off cleanly. The gap between a deal closing and a person being staffed runs three to six weeks that neither the forecast nor the staffing model covers on its own, and that gap is where most capacity-planning failures live.
A system that merges the two closes this gap, because the inputs finally live in one place. But the gap is part data and part decision, and merging the data only solves the first half. Someone still has to take a weighted pipeline number and a bench number and decide who would be staffed if each deal closed: where is each project, what is its scope, who can deliver it, are they available, should the kickoff move to free someone up. That assembly work, gathering the inputs and laying out the options, is what eats the three days. Once the pipeline and the staffing model sit in one system, that legwork is mechanical enough that AI reading the staffing model against the live pipeline can do it and propose the call, leaving the human with the judgment rather than the assembly. Servantium is built around exactly this merge.
The hours you schedule versus the hours you bill
There are two capacity numbers, and most plans show only the first. Scheduled capacity is the person-hours committed in the scheduler. Billable capacity is what actually converts to billable work, once you subtract internal meetings, training, business development, and the unbilled hours that exist in every engagement. The gap is typically 12 to 20 percentage points: a team scheduled to 89% is usually billing closer to 71%. A plan that shows only the scheduled number overcommits the team, and overcommitment is how the work gets underpriced, most painfully on anything fixed-fee.
Industry data confirms the stakes: average billable utilization across professional services firms fell to 66.4% in 2025, down from 68.9% in 2024, while EBITDA dropped to 9.8% from 15.4% the prior year, according to the 2025 Professional Services Maturity Benchmark by Kantata. Teams that measure only scheduled utilization cannot see the billable number falling, and cannot act on it.
Capacity planning that does not show both numbers is performing capacity planning, not doing it. The decomposition has to be visible in the model. If the model shows only scheduled hours, it is flattering you.
The $50,000 project that ran for five years
I avoid fixed-fee work when I can. The time it taught me the most, I had no choice. I inherited an account where the same $50,000 project had been open for five years, and I was handed the latest copy of the same goal everyone before me had been handed: finally get them live, this time under my leadership.
The account had referred us a long list of customers, and it knew it. That leverage meant the project went out of scope on day one and no one ever reeled it back in. Integrations, migrations, customizations. By the time it reached me, the product they were running looked nothing like what had been promised, sold, or supported. Every path forward was bad. Restart it, or keep trudging.
We restarted, with about five people on it. Within a month, once the real effort was finally visible, the organization did the thing that doomed it: it looked at the cost and scaled the team back to two. Nobody had ever put an honest number on what the work required, so when the number finally showed up, it was easier to cut staffing than to believe it. Two people could not close a five-year gap. The project never went live.
That is the part worth sitting with. The effort was not unknowable, it was unmeasured, and an unmeasured number is easy to wish away. The gap between what a project needs and what it is staffed for stays invisible until someone forces the calculation into the room where the staffing decision actually gets made. If the only number on the table is the one that keeps the quarter looking fine, that is the number you will staff to, and you will lose.
The Friday capacity check
Three people meet for 30 minutes: the pipeline owner, the staffing owner, and one senior practice lead who breaks ties. One screen shows four panels: weighted pipeline for the next eight weeks, bench by skill for the same period, current overallocations by named person, and last week’s open decisions. They walk left to right, the pipeline owner names each near-closing deal, the staffing owner says yes or no, the practice lead adjudicates. The output is three to seven recorded decisions. That is the ritual.
Whether you hit 30 minutes or 90 depends entirely on data hygiene. If the pipeline is honestly weighted and the bench numbers are clean, the conversation is fast. If you are sorting out stale CRM stages and debating whether Alex is really on PTO that week, the ritual will expand to fill the data gap. The meeting is not the hard part. Clean data is the hard part.
Who attends. Three people: the person who owns pipeline forecasts, the person who owns delivery staffing, and a senior practice lead who can adjudicate trade-offs. If a CFO wants to be in the room, fine, but the meeting has three working voices.
What is on the screen. One view, four panels. Top-left: weighted pipeline for the next eight weeks, deal-by-deal, with weighted person-week demand. Top-right: bench by skill for the same eight weeks, with absences subtracted. Bottom-left: overallocations flagged, by named person. Bottom-right: decisions pending from last week’s meeting, with status.
What happens. Walk left to right. The pipeline owner names each at-risk-of-closing deal and what it would need. The staffing owner says yes, no, or that it needs more discussion. The practice lead breaks ties. Decisions get recorded with rationale.
What does not happen. Slides. Pre-reads. Dashboards as ceremonial objects. The whole thing runs on one screen and produces three to seven decisions a week. The decisions are the output.
Here is the shape of a week-one model. Four roles, one week.
| Role | Available hrs | Committed hrs | Free hrs | Notes |
|---|---|---|---|---|
| Partner | 20 | 14 | 6 | 3 hrs BD, 1 hr admin |
| Senior consultant | 36 | 32 | 4 | Soft commit on Deal B |
| Consultant | 40 | 38 | 2 | At risk of overallocation if Deal B closes |
| Analyst | 40 | 24 | 16 | Available for Deal C ramp |
Available hours are scheduled hours minus PTO, training, and a 15% overhead buffer for internal work. Free hours are what is genuinely reclaimable, not what the scheduler shows as unassigned.
Run this for your three busiest roles across the next eight weeks. If Senior Consultant free hours drop below four in any two consecutive weeks, you have a constraint that will surface as a staffing scramble when the next deal closes. The table is not the model. The table tells you which conversation to have on Friday.
Three horizons, not one
Run three separate models at three separate cadences: operational (0-8 weeks, named person, weekly), tactical (2-6 months, skill category, monthly), and strategic (6-24 months, practice headcount, quarterly). The most common failure is running the tactical model in the operational slot. At a 50-person team, a senior consultant going on leave, a partner canceling vacation, or a soft-commit shifting can each change whether you can take a deal within a week. Monthly reviews cannot track that.
| Horizon | Range | Granularity | Cadence | Primary question |
|---|---|---|---|---|
| Operational | 0-8 weeks | Named person, weekly | Friday 30 | Can we take this deal if it closes Thursday? |
| Tactical | 2-6 months | Skill category, monthly | Monthly brief | Are we hiring ahead of Q3 demand? |
| Strategic | 6-24 months | Practice headcount, quarterly | Quarterly plan | What does the team look like in 18 months if the pipeline holds? |
The Friday 30 is not a substitute for the quarterly hiring plan. Run weekly. Brief monthly. Plan quarterly. Keep the three horizons in separate rooms.
Treat the named-person forecast as a record, not a derived view
A named-person forecast tracks each individual’s allocation to closed work, soft commits, and weighted pipeline, week by week, for the next 12 weeks. It is updated continuously by the people who own staffing. It is the system of record for whether you can take a deal. The derived version (headcount times target utilization) is right for board reporting. It is wrong for Tuesday afternoon when a deal closes early and you need to know if Jamie can hold June 8 to 19.
The single move that separates firms with working capacity planning from firms with broken capacity planning is treating the named-person forecast as a record the system maintains directly, not as a view derived after the fact.
Most tools default to the derived version: 12 senior consultants, average utilization target 78%, monthly billable capacity equals 12 x 160 x 0.78 = 1,498 hours. That number is right for board reporting. It is wrong for the operating question. The CFO and the board run on the derived version. The Friday 30 runs on the first-class version. If your firm only has the derived version, capacity planning produces the wrong answer to the question that matters on a Tuesday afternoon when a deal closes early.
When a partner wants to pull a consultant
Three situations make an override legitimate: genuine revenue acceleration (the work is in the weighted pipeline and capturing it now is worth the disruption), a client save (a material relationship is at risk and this consultant is the specific fix), or a board-priority engagement flagged in advance. Every other override pattern, including moving a consultant to show commitment on a deal not yet in the weighted pipeline, defaults to no. The decision sits with the VP of Professional Services or COO, not the requesting partner. Every override gets logged.
Every capacity planning ritual breaks at the same point. A partner walks in on Friday morning with a client they need to keep happy and a senior consultant they want pulled off another engagement. The capacity model says no. The partner says yes. The room goes quiet and someone blinks.
Without a rule decided before the moment arrives, the ritual collapses. The plan gets overridden, the decision does not get recorded, and two weeks later the same fight happens again with no institutional memory of how it went last time.
All three legitimate situations share one structural feature: the override serves the firm’s declared pipeline, not the partner’s personal book of business.
The most common pattern that defaults to no is jam-tomorrow billing: a partner has a verbal conversation with a contact, believes the work will materialize, and wants to move a senior consultant now to show commitment. The deal is not in the weighted pipeline. The consultant pulled is unavailable for committed work, and the firm has moved capacity against a guess. Two other patterns also default to no: scope creep dressed up as a new deal to get resource priority, and personal pull, where the partner wants a specific consultant because they are comfortable together, not because the work requires their skills.
The override decision sits with the VP of Professional Services or the COO. Neither the requesting partner nor the staffing owner resolves it unilaterally. Every override, approved or denied, goes into the capacity log: date, partner who requested it, consultant affected, rationale. The log gets reviewed at the monthly brief. One partner generating repeated denied requests is a management conversation, not a Friday morning argument.
Frequently asked questions
-
Calibrate close rates by stage from your own 24-month win history. Typical ranges: 8% to 15% at qualified, 25% to 40% at proposal, 55% to 75% at verbal, 85% to 95% at signed. Multiply each deal by the stage rate and sum across the pipeline. The pipeline is unpredictable per deal but stable per stage. Run this calculation in the Friday 30, not once a quarter.
-
Hire to average. Bridge to peak with contractors and partners. Hiring to peak guarantees bench during troughs and destroys margin. The mistake firms make is bridging to peak with overtime instead of flexible capacity, which works for two quarters and then burns out the team. Build contractor relationships before you need them, not during the scramble.
-
It depends on billing model and role mix. Delivery consultants on fixed-fee retainers can run at 75% to 85%. Senior consultants who carry proposal and mentoring obligations should sit closer to 65% to 70%. Partners building new books of business often run 30% to 50% billable. The number that matters is billable utilization, not scheduled, since the gap between them runs 12 to 20 points. Set targets against the billable number and the scheduled number stops flattering the plan.
-
Three horizons, kept separate. Operational: eight weeks, weekly granularity, in the Friday 30. Tactical: six months, monthly granularity, reviewed monthly. Strategic: 18 to 24 months, quarterly granularity, reviewed with the hiring plan. The operational rhythm should not get polluted with strategic uncertainty, and vice versa.
-
A named-person view covering the next eight weeks, updated every Friday before the check-in. Three columns per person: committed hours (closed work), soft-committed hours (weighted pipeline over 60%), free hours. The conversation is the decision. The view is the shared screen that makes the conversation fast. A spreadsheet works to start, but it breaks the moment the pipeline and the bench need to update each other in real time. That is the point where merging the two into one system, with AI reading the staffing model against the live pipeline, earns its keep.