P17 min readUpdated

Mobile QA beyond “run it on BrowserStack”

Device farms are useful. They are not a strategy. What a farm can tell you, what only a phone in the hand shows, and what I put in a mobile test plan.

Physical deviceBrowserStack matrixWearable journey

“Tested on mobile” usually means an automated suite ran on a device farm and came back green. Nobody held a phone. Then a user installs the update over a version from months ago, standing in a car park with one bar of signal, and meets a permission prompt they refuse, a keyboard that covers the pay button and a session that expired while the phone was in their pocket.

Mobile app testing has been a thread through my work since I started out testing mobile games: native iOS and Android on a smart-city parking platform, mobile banking, mobile web and, most recently, apps paired with wearables. BrowserStack is part of how I work and I would not give it up. But a device farm is a tool, and a list of farm runs is not a mobile test plan.

What a device farm is good for

A farm such as BrowserStack gives you breadth and speed. In an afternoon I can see one build on more combinations of OS version, screen size and manufacturer than any team keeps in a drawer. An automated regression suite can run on many of them in parallel, with video and logs for every failure. It is also the quickest way to see a defect reported on a phone model nobody on the team owns.

So the farm answers one real question well: does this build run across the range? For regression on a stable core, that is most of what I need.

Device farm vs physical device: what each one shows

The limits come from what a farm is. The phone sits in a rack in a data centre, on a good connection. I reach it through a video stream and drive it with a mouse. It is wiped between sessions, so every test starts on a device with no history. Hardware the app depends on, such as the camera, biometrics or Bluetooth, is simulated, injected or out of reach.

What I want to knowDevice farmPhone in the hand
Does the layout hold across OS versions and screen sizes?Strong: many combinations, quicklyWeak: only the few devices you own
What happens when the app updates over months of real use?Weak: each session starts cleanStrong, if one device is never wiped
What happens when the signal drops or changes mid-journey?A throttled network profile at bestStrong: walk to the lift or the car park
Do the camera, biometrics and Bluetooth work with real hardware?Simulated or limitedStrong
Can it be used one-handed, outdoors, with the on-screen keyboard?Cannot tellStrong

The right-hand column is where users live. A phone in the hand catches install friction, permission prompts, keyboard behaviour and the awkward paths people hit between buildings and lifts, where the connection moves from wifi to mobile data and back. A throttled profile on a farm slows the network down evenly. It does not reproduce the moment the connection changes halfway through a payment, when the server has taken the request and the app never hears the answer.

What a mobile test plan should cover

The point of a mobile plan is not collecting devices. It is mapping the journeys that cross app, API and hardware, then deciding where each one has to be exercised. Mine usually covers six things.

  • Install and upgrade. A fresh install, and an upgrade over the previous store version with a signed-in user and real data on the device. I check what survives: the session, saved preferences, anything stored locally, permissions already granted.
  • Poor network and offline, where the journey meets them. What the app shows when a request hangs, and whether a retry repeats the action. A payment that is retried should still be taken once.
  • Push and deep links. Notifications with the app open, in the background and closed. A tap lands on the right screen whether or not the user is signed in. A link in an email opens the app when it is installed and something sensible when it is not.
  • Interruptions. A call, a locked screen, a switch to another app and back, a return after an hour. The journey resumes or says clearly that it cannot.
  • Permissions. The first ask, the refusal and the way back, for camera, location and notifications.
  • A small real matrix. A short list of devices and OS versions, agreed with Product, with the critical journeys walked by hand on each.

When the product reaches beyond the phone, the chain gets longer. Wearables add pairing, health permissions and background sync, and none of those can be judged from a screenshot.

How many physical devices you need for mobile testing

Fewer than people fear. I choose from the product’s own analytics and from who the users are, and I cover classes of device before brands: a current iPhone, and an older, smaller one on the oldest iOS the app supports; a current Android flagship; and a mid-range Android from a manufacturer that customises the system heavily, because aggressive battery management is where background work and notifications tend to die. A tablet joins the list only if people use one.

One of those phones is never wiped. It carries an account that has existed for months, takes every build as an upgrade and keeps whatever permissions it was given. It is the nearest thing I have to a real user’s device.

The farm takes breadth and the automated regression. The physical devices take a few deliberate sessions per release, on the journeys that changed and the ones that carry money or trust. The release readiness note then says which was which: what ran on the farm, what was walked by hand on which phones, and which journeys nobody held a device for.

The objection: the farm has real devices too

It does, and that is the best argument for a farm over an emulator. What is missing is everything around the hardware: a hand, a pocket, a weak signal, a history. The other objection is budget. Four or five phones are a small purchase beside a release that leaves paying users stuck at sign-in.

Something else happens once testers do hold the phone: the sessions turn up behaviour that meets the acceptance criteria and still fights the user. The keyboard hides the field being typed into. The error after a dropped connection says “try again” without saying whether the payment went through. None of it fails a test case. I raise these as improvement stories or low-priority defects, written as user impact Product can weigh. That is not scope creep. It is cheaper than meeting the same friction in UAT or in production.

If your mobile QA is only a BrowserStack checklist, you are testing screenshots of confidence. Pair the farm’s coverage with a few deliberate physical-device sessions and the readiness note can say something a green run cannot: a person completed this journey, on this phone, in the conditions your users will be in.

Questions I get asked

Is a device farm enough for mobile app testing?

For breadth, yes. A farm shows quickly whether a build installs, lays out correctly and completes scripted journeys across many OS versions and screen sizes. It is not enough for behaviour that depends on a real hand and real conditions: upgrades over months of use, a signal that drops mid-payment, interruptions, the camera, biometrics and Bluetooth. Pair it with a small set of physical devices and record which checks ran where.

How many physical devices do you need for mobile testing?

Usually four or five, chosen from your own analytics: a current iPhone and an older one on the oldest iOS you support, a current Android flagship and a mid-range Android from a manufacturer that customises the system, plus a tablet if your users have one. Keep one device that is never wiped, so every build is tested as an upgrade on an account with history.

What is the difference between an emulator, a device farm and a real device?

An emulator or simulator is software on a computer that imitates a phone. A device farm is a set of genuine phones in a data centre that you control remotely through a browser. A physical device is one you hold. The first two are good for breadth and fast feedback. Only the third shows how the app behaves in a hand, on a real network, with real sensors and a history of use.

Filed under Mobile and wearable QA. Terms: Device farm, Exploratory testing.