Every ERP proposal has a line called “data migration”. In the ones I see from other suppliers it is usually somewhere between eight and fifteen days, sits near the bottom of the schedule, and is the single most reliable predictor of whether the project will over-run.
It is not that eight days is always wrong. It is that nobody knows whether it is wrong until somebody has read the data, and on most projects nobody has read the data at the point the number is written down.
Where the days actually go
Below is the breakdown from a genuine mid-market ERP migration: a manufacturer, two sites, roughly 4,000 items, three years of transaction history, moving off a legacy ERP with a surrounding estate of spreadsheets. Forty-one days in total, against an initial estimate of twenty-two.
| Activity | Estimated | Actual | Why it moved |
|---|---|---|---|
| Source profiling | 3 | 5 | Four undisclosed spreadsheet sources |
| Mapping design and sign-off | 4 | 7 | Three fields meant more than one thing |
| Build — extract and transform | 8 | 9 | Broadly as expected |
| Dry runs (4, not 2) | 4 | 10 | Two runs failed to reconcile |
| Reconciliation reporting | 2 | 6 | Finance needed it cut three more ways |
| Cutover and post-load checks | 1 | 4 | Weekend cutover, two people |
| Total consultant days | 22 | 41 |
At our data migration engineer rate that is a difference of just over £18,000 — on a project where the same client spent three weeks negotiating a £6,000 reduction in the implementation estimate. The negotiation was on the number they could see.
The three things that always cost more than expected
1. Sources nobody mentioned
Ask an organisation how many systems hold data going into the new ERP and you will get a number. It is always low, and not because anyone is being evasive. Departmental spreadsheets are not thought of as systems by the people who maintain them, right up until you try to go live without the pricing matrix that lives in one.
On the project above, four sources surfaced during profiling that were not in the original scope, including a supplier price list maintained by one person in purchasing that turned out to be the authoritative source for about 900 items.
2. Fields that mean more than one thing
This is the expensive one, and it is a business problem wearing a technical costume.
In a system that has run for fifteen years, at least one field will have been repurposed. A customer reference column that holds a genuine reference for records before 2019 and a route code after it. A status field with eleven values, four of which are no longer used and two of which mean the same thing for historical reasons nobody can now recall.
You cannot map that field until somebody decides what it should mean in the new system, and that decision is not ours to make. It needs the person who knows the business, and their availability is the actual constraint. Three such fields cost three days here — not of engineering, but of waiting, meeting, and re-working the mapping afterwards.
3. Reconciliation that has to satisfy finance
Everyone budgets for a reconciliation report. Few budget for the fact that the first one will be rejected.
A migration is not finished when the data lands. It is finished when someone in finance agrees the totals — and they will want them cut by period, by account, by cost centre, and probably one more way you did not anticipate. That is entirely reasonable. They are the ones signing off that the opening balance is correct, and “the row counts match” is not an answer to that question.
How we estimate it instead
We do not quote a migration before profiling the source. Not as a commercial posture — we simply do not have the information, and neither does anyone who does quote it.
The sequence we use is: profile first, as a small fixed-price piece of three to five days; then estimate the migration against the profiling report; then build. The profiling report is yours whether or not you proceed with us, and several clients have taken it to a competitor. That is fine. It still produced a better-informed decision than the alternative.
What the profiling report contains:
- Every source, including the ones discovered rather than declared
- Row counts, distinct values, null rates and orphan records per table
- Fields whose meaning has drifted, with the evidence
- Duplicate and near-duplicate analysis on master data
- An estimate with a range, and the specific things that would move it
The uncomfortable part
Some of what profiling finds is not a migration problem. It is a data quality problem that predates the project by years, and it becomes visible because we counted something nobody had counted before.
We do not fix that. We cannot: correcting nine hundred duplicated supplier records requires someone who knows which is the real one, and that is not a decision an external engineer should be making. What we do is quantify it precisely, so the effort can be planned rather than discovered on cutover weekend.
Clients occasionally find this frustrating, and I understand why — you engaged us to move the data and we have handed you homework. But the alternative is not that the problem goes away. The alternative is that it moves into your new system, where it will be more expensive and more visible.
The best migration I worked on last year over-ran its original estimate by 40% and went live with a clean reconciliation on the first attempt at cutover. The client considered it a success, because we had told them in week two exactly why it would over-run and by roughly how much. Bad news early is just information. Bad news late is a project problem.