← All Articles

How to Start Process Mining: A Step-by-Step Assessment Guide

Plan a process mining assessment from the business question to event-log validation, variant analysis and improvement decisions, with a sample CSV dataset.

Discovery guide 5 of 6

Useful for
Business analysts, data owners, process improvement leads and automation architects
What you will take away
Prepare a defensible mining assessment with defined events, quality checks, reproducible findings and an improvement hypothesis.

Start process mining with a defined business question and a check that the available event history can answer it. Agree the process boundary and case definition, validate an extract against source records, analyze a representative population, and review the findings with practitioners before recommending a change.

This is a practical assessment guide, not a product installation tutorial. It includes a synthetic invoice dispute event log that you can inspect in a spreadsheet or adapt to a supported mining tool.

When is process mining a useful starting point?

Assess it when you need to understand recorded process variants, repeated activities, handoffs or delays across a population of cases. It may help investigate why the same business process behaves differently across sites, document types or channels.

A one-off, well-understood task may need only targeted observation and records. If the system stores only current status or lacks the history relevant to your question, start by resolving that evidence gap. The workshop versus process mining guide explains when a lighter approach or a combination is appropriate.

Step 1: define a question that can change a decision

Replace “find savings with process mining” with a question such as: “which invoice dispute categories repeatedly request delivery evidence, and where is the associated delay?” Name the process owner, intended decision and measure.

Write the trigger and completed outcome. Agree whether you are analyzing a dispute, an invoice, an invoice line or another business object. One invoice may have several disputes or shipment references. Joining them without a clear relationship can duplicate events or attribute work to the wrong case.

Use an initial SIPOC workshop to settle scope and exclusions. Agree which events should exist and where work may occur outside the system.

Step 2: check data access and event definitions

For a conventional case-based event log, identify a case ID, activity name and event timestamp. An activity log may additionally provide start and end timestamps. Microsoft distinguishes these formats in its data requirements. Microsoft: Prepare processes and data.

The following is a proposed analysis dictionary for our teaching example. Product-specific field names and mappings vary.

Field Example Question to resolve
CaseId D001 Does this uniquely identify one dispute within the selected scope?
Activity Evidence attached What business event creates this record?
Timestamp 2026-09-01T09:10:00Z Is this the occurrence time, a later update or an ingestion time?
BusinessUnit North Is the attribute stable, or can it change during the case?
DisputeType Delivery evidence Who defines the category, and are historical values consistent?

Approve access and the intended use with the relevant data owners. Limit fields to the purpose, use appropriate permissions and handle personal or commercially sensitive data according to the organization’s controls. Desktop task capture requires a separate scope and review if it is introduced.

Keep source-to-analysis mappings and extraction logic. An analyst should be able to explain what each event represents without relying on a dashboard label alone.

Step 3: validate a small extract before scaling

Trace selected cases from the extract back to the source application with a practitioner or application owner. Include an ordinary case, a repeated event and an unfinished case.

Check for:

  • Missing or reused identifiers, including collisions across business units.
  • Duplicate events introduced by extraction or one-to-many joins.
  • Time zones, inconsistent formats, impossible sequences and tied timestamps.
  • Batch updates that appear to be business activity but represent system maintenance.
  • Missing history, deleted records and events outside the selected period.
  • Open cases and cases that began before the observation window.
  • Changes in category definitions, logging behavior or source-system versions.

Document a disposition for each issue: correct the mapping, retain with a qualification, exclude with a reason or pause the analysis. Keep counts before and after transformations. Do not silently drop the difficult cases to produce a cleaner process map.

Step 4: use a sample event log to understand the analysis

Download the synthetic invoice dispute event log (CSV). It contains 12 events across three invented cases. It is deliberately too small for a business conclusion; use it to learn the mechanics. Treat 3 September 2026 at 12:00 UTC as the observation cutoff.

Case Selected events What the example illustrates
D001 Registered 1 September 09:00; closed 10:00 the same day A completed path lasting one elapsed hour
D002 Registered 1 September 09:00; evidence requested twice; closed 2 September 10:00 Repeated requests and 25 elapsed hours
D003 Registered 2 September 09:00; evidence requested; no closure by the cutoff An open case with 27 hours of age at the cutoff

In a spreadsheet, sort by case and timestamp, read the ordered activity sequence and compare the first event with the defined closing event for completed cases. Keep open-case age in a separate measure. Do not manufacture a closure timestamp for D003.

The two completed cases have a mean elapsed duration of 13 hours: (1 + 25) ÷ 2. That does not mean 13 hours of staff effort were consumed. Nor does this tiny dataset prove that a repeated evidence request caused a 24-hour increase. Those are questions for additional evidence and practitioner review.

Step 5: analyze variants, rework and time with clear denominators

Begin with population and data-quality counts before performance charts. Report how many cases and events are included, the observation period, exclusions and open-case treatment.

Then examine:

Analysis Definition to agree Decision it may support
Process variants Ordered paths through the included activity set Which variants need different design or further investigation?
Repeat activity A specified activity occurring more than once in a case Is repetition expected, administrative or avoidable?
Elapsed duration Time between agreed starting and closing events Where should delay be investigated?
Handoff gap Time between two meaningful recorded events Which cases should be reviewed for waiting or unrecorded work?
Completion and open-case age Closed outcomes and age of unfinished cases, reported separately Is unfinished work accumulating or being hidden by completed-case averages?

A proposed repeat-request rate is cases with more than one evidence request divided by cases in the eligible population, with the period and eligibility rules stated. Do not divide repeat events by cases and label the result a case rate.

Segment results by relevant, supported attributes. Compare distributions and outliers alongside averages. Check whether a pattern reflects a different case mix, volume change or logging convention before comparing teams.

Step 6: validate explanations with business users

Take examples behind each finding to the people performing the work. Ask what happened, which information was unavailable, what authority was required and whether the proposed event interpretation matches the source.

For the synthetic D002 case, two requests could reflect missing shipment data, a legitimate request for additional evidence or duplicate logging. A workshop can help identify which explanation to test, but the sample should still be checked against records.

Measure active handling separately where labor effort matters. An interval between events is not automatically productive work or a recoverable saving. Use approved observation or suitable activity evidence to establish the effort baseline.

Step 7: turn findings into an improvement hypothesis

Write each hypothesis as a specific change to a defined population, with an owner, prerequisite and measure. For example: “provide a validated shipment identifier at dispute intake, then assess whether repeat evidence requests decrease for the supported category.”

Compare changes to input quality, rules, ownership, application configuration and automation. A retrieval workflow might use approved APIs; a permitted portal task might need RPA; varied notices might justify evaluated document extraction. A mining tool identifies evidence for the decision, not the execution stack by itself.

Agree a bounded trial, business acceptance scenarios and a comparison basis. Track both expected improvement and extra work or failure modes introduced by the change.

Which process mining tools should you evaluate?

Start with existing access and capabilities. A spreadsheet or SQL analysis can help validate event meaning and scope before a platform trial. For a broader assessment, Power Automate Process Mining and UiPath Process Mining are examples to consider, not a mandatory shortlist or partnership claim. Microsoft overview, UiPath overview.

Evaluate whether the proposed tool can support your case relationships, event definitions, data volume, access controls, analysis and refresh needs. Ask who will maintain extraction and review the findings. Demonstrate the important questions with representative data before selecting the platform.

This guide does not estimate product prices or implementation fees. A suitable tool is one the organization can operate and use to make better decisions, not simply one that renders an attractive diagram.

What should the assessment deliver?

Expect an agreed question and case boundary, source and event dictionary, documented transformations, data-quality findings, analysis definitions, validated observations and prioritized improvement hypotheses. Record what the data cannot establish.

Download the process mining readiness worksheet (CSV) to capture the assessment evidence and next decisions. Use the benefit estimation guide before converting process findings into calculator inputs or a savings claim.

For an ongoing initiative, connect accepted improvements to the existing roadmap and benefit owners. The existing-program discovery guide explains how to avoid overlapping work.

Common questions about getting started with process mining

Can a current-status export be used as an event log?

It may describe the current population but generally cannot reconstruct activities that were never retained. Establish whether historical records or reliable event timestamps exist before claiming a view of past paths.

How much data is enough?

Enough to cover the variants and operating conditions relevant to the question, with known limitations. There is no universal sample size in this guide. A small extract tests data readiness; it does not automatically establish a population-wide pattern.

Can we calculate staff savings directly from process duration?

No. Elapsed duration includes waiting and may miss work outside the system. Establish active handling and residual effort separately, then validate the proposed change and realization plan.

Have a Process Like This?

Discuss your workflow, the evidence you have and the questions to resolve before choosing an implementation approach.

Step 01

Understand the Process

Discuss one workflow and its pain points.

Step 02

Identify Constraints

Discuss systems, data and exceptions to assess.

Step 03

Explore Potential Value

Identify the effort, volumes and costs to validate.

Step 04

Agree the Next Step

Decide whether a deeper assessment would help.

Book a Strategy Session
✓ Zero obligation ✓ Start with one workflow ✓ Talk directly with an automation practitioner