Work

Mobile Test Automation Practice 2025

A team shipping a revenue-critical Flutter app had no automated tests. Regression was done by hand, which meant only the headline features got checked, the check cost days, and the sign-off was somebody saying it looked fine. I built the practice that replaced that, and the method behind it.

Akkodis blog article LinkedIn post

01

What manual regression was costing

Time

A full pass took a QA day or more, and every retest spent it again.

Narrow coverage

Only the headline features were checked, because that is all a person can reach by hand.

Late detection

Defects surfaced right before a release, or after it.

No evidence

Sign-off was a verbal yes, with nothing to audit afterwards.

02

One run, start to verdict

One command

Kick it off and walk away

No babysitting and no terminal skills required. A picker chooses the areas, the user profile, the language and the tag filter, so QA runs it without touching a command line.

Unattended

It drives the real app on a device

Taps, types and scrolls a real build. Black box, no stubbing, hard timeouts so a hung step cannot hold the run open.

Every checkpoint

Evidence captured as it goes

A screenshot at each checkpoint and a recording of the whole run, so a failure arrives with the exact step and a picture of the device at that moment.

At the end

The report grades itself

A self-contained HTML file with a computed GO or NO-GO verdict, per-flow results, durations and video timestamps. Never hand-written.

Then again

A second program checks the first

An independent verifier reads the same evidence and recomputes every number, so the grade is not the opinion of the thing being graded.

90+

journeys automated, from about 10 by hand

~12 min

one full unattended run

2 languages

English and Arabic, RTL included

03

What one run replaces

The report computes this for itself on every run, against the manual baseline it was built to replace.

Manual baseline

157 UAT cases at about 5 minutes each, roughly 785 minutes, or 13 hours of skilled QA time. Headline features only, and no evidence at the end of it.

This run

12 minutes, unattended, with every scenario screenshot-backed.

Saved per run

About 773 minutes, close to 13 hours, every single time it runs.

And repeatable

At a cadence manual regression cannot reach, because nobody has to be there.

04

How it is put together

The four moving parts of the test harness, left to right. One: flows, YAML journeys with one file per scenario, keyed to stable widget IDs. Two: the runner, which executes one flow, the full suite, or two lanes in parallel. Three: the report and verifier, a self-grading HTML report whose numbers a separate verifier recomputes. Four: the CI gate, a device-free lint on every push plus cloud runs on both platforms.

Everything specific to a project lives in one config file, so the engine is not tied to the app it was built against. Pointing it at a second Flutter app is a new config and that app's own flows. It has been split into its own repository and run unchanged.

05

The unlock was in the app, not the tool

A Flutter app renders to a canvas. Without semantics, an automation tool has nothing to address and can only tap coordinates, which stop being true the first time the layout changes. One semantics ID per widget, shipped with the screen, is what makes the app testable at all. It is a one-line change in a pull request, and it is also an accessibility win, which is the argument that got it adopted.

Two ways an automation tool sees the same Flutter screen. On the left, with no semantics IDs, the screen is one opaque texture with nothing addressable, so a tap can only be aimed at a coordinate. On the right, with semantics IDs shipped, the same screen is an accessibility tree of named nodes: search field, product card, add to cart button and checkout submit.

06

Why Maestro and not Appium

DimensionAppiumMaestro
Sees Flutter widgetsYes, once the app exposes semantics IDs. Before that, an opaque canvas.Yes, through the same accessibility tree.
Flutter supportBlack box through UiAutomator2 and XCUITest. The Flutter driver bridge is deprecated.First class, with no driver to install.
FlakinessManual waits are typical, and there is a driver stack to keep working.Auto-waits, one code path.

The honest version of this table: the real unlock was the semantics IDs, and those would have helped Appium too. Maestro won on lighter setup, lower flakiness, and because the report and verifier were built around it.

07

Eight gates the report grades itself against

A verdict is only worth something if the thresholds were set before the run rather than after it. These are the thresholds, and every report computes itself against them.

GateWhat it measuresGreenAmberRed
P0 pass rateMust-pass scenarios that passed this run>= 95%90 to 94%< 90%
P0 coverageShare of must-pass scenarios actually automated>= 80%70 to 79%< 70%
Overall pass rateEvery flow that passed, across every area>= 90%80 to 89%< 80%
FlakinessFlows that pass and fail inconsistently across runs<= 5%6 to 10%> 10%
Suite freshnessDays since the last fully green full run<= 78 to 14> 14
Last-run recencyDays since the suite ran at all<= 34 to 7> 7
Blocker ageDays the oldest open blocker has stayed open<= 1415 to 30> 30
Stability streakConsecutive runs with the same green result>= 53 to 4< 3

08

Breaking it on purpose

A journey counts as covered only when its unhappy paths are automated too. One epic does nothing but misuse the app.

×

Rapid taps, double submits, and the race conditions they expose

×

A wrong one-time code, repeated: the screen must not advance

×

Backgrounding and relaunching mid-flow: state has to survive

×

Hostile text in fields: very long, emoji, right-to-left

×

No connection at all: fail gracefully, do not crash

This is the part that has already paid for itself. Driving the app the way a confused or impatient person would surfaced a reproducible crash on both platforms, found from the evidence trail, investigated against logs, and re-verified before it was reported.

09

Who owns it after me

A practice that only works while one person is present is not a practice. The roles are written down, and so is the method.

QA

Runs the picker, triages a failure into app bug, test bug or data, and writes flows for their own cases.

Developers

Ship one semantics ID with every new screen. Free in the pull request, and an accessibility win.

Maintainer

The engine, the reliability patterns, and the catalogue of failure modes.

MaestroFlutterPythonYAML flowsGitHub ActionsAndroidiOS Simulator

10

Written up

Automating Mobile App Regression, on the Akkodis blog

A Mobile Test Automation Practice, from Zero, the methodology