The short answer: there are five kinds of software testing partner — functional (manual) QA, test automation, API and integration testing, performance or security specialists, and a managed QA service — and buying the wrong one is the mistake that ruins most engagements. Decide which you need first, then evaluate candidates on three things: whether the tests they write are readable and runnable by your own engineers, how honestly they report quality, and what happens when they find nothing at all. Price is the last criterion, not the first.
Why testing partners are unusually hard to compare
Every testing vendor's website says the same four things: rigorous testing, faster releases, higher coverage, quality at scale. None of it is evidence, and it is not meant to be. It is the entry fee for the category.
What makes this purchase different from buying development capacity is that the deliverable is mostly invisible. You can look at a shipped feature and judge it. You cannot look at a test suite and see whether it would have caught the failure that takes production down at 2am next quarter. Coverage percentages are self-reported, and a thousand tests that assert nothing still produce a green check mark.
A useful reframe: you are not buying test execution. You are buying judgement about what is worth testing, plus the discipline to keep that judgement current as the product changes. Execution is the part that automates well. Judgement is the part that has to be bought.
The five kinds of testing partner
| Partner type | What they own | Typical trigger | What to watch |
|---|---|---|---|
| Functional / manual QA | Exploratory and scripted testing of real user journeys | New product, or no defined test strategy yet | Coverage claims with no traceability back to requirements |
| Test automation | Building and maintaining an automated suite inside your CI | Regression runs are blocking releases | A suite only the vendor can run or repair |
| API, integration and contract testing | Service-level tests, contracts and third-party integrations | Microservices, partner APIs, data pipelines | So much mocking that nothing real is ever exercised |
| Performance, security or accessibility specialists | Load, penetration and WCAG testing | A launch, a compliance deadline, a scale event | A one-off report with no retest included |
| Managed QA service | The whole QA function, including strategy and leadership | No QA leadership in-house | You outsourced the thinking as well as the clicking |
Most organisations end up needing two of these, not one: an engineer or small team for ongoing product work, and a specialist engagement for a specific event. Be explicit about which problem each one solves. A specialist firm hired for continuous regression work will be expensive and disengaged, and a generalist partner asked to sign off a penetration test will not be able to.
The five criteria that actually predict success
1. Tests you can read, run and maintain without them
The best single predictor of a working engagement: ask to see a sample of test code they have written, and ask how long it takes a new engineer on your side to run it. Good automation is boring. It is named clearly, driven by the business language of the product, and runnable with one command in your own CI.
If the tests live in a vendor's platform, run from a vendor's dashboard, or read like transcribed click coordinates, you are renting quality rather than building it. That is a legitimate choice — but price it as renting, and assume the asset leaves with the vendor.
2. How they decide what not to test
Exhaustive testing is not a strategy; it is a budget with no end. A mature partner will push back on a list of everything you want covered and propose a risk-weighted subset instead: the flows that touch money, personal data or a regulatory obligation first, the cosmetic paths last.
A partner who agrees to test everything you name, at the price you name, has just told you they will not exercise judgement on your behalf.
3. What their defect report actually contains
Ask for a redacted example of a real defect ticket. It should carry the environment, the exact steps, the expected and actual behaviour, evidence, and a severity with a stated rationale. Most importantly, it should reproduce for someone who did not write it.
What you are looking for is triage discipline: not every cosmetic issue escalated to release-blocking, and not every data-corrupting bug filed as low priority because it is awkward to fix late in the cycle.
4. Whether they will tell you the release is not ready
This is the test of the commercial relationship. A partner whose revenue depends on billable test cycles has an incentive to find more work. A partner worth keeping is willing to say that the build is acceptable and that further testing is not the best use of your budget this month.
Ask directly, in the first meeting: when did you last advise a client to ship when they expected you to find more work? Then watch the answer.
5. How knowledge transfers when the engagement ends
Testing engagements end — for budget reasons, for a reorganisation, or because the work is finished. Ask what the exit looks like: test suites in your repository, documentation of the strategy and the known gaps, a handover session with the engineers who will maintain it, and a written list of the risks you are accepting by stopping.
A partner who only discusses the exit after a resignation has landed is telling you they never planned for one.
How testing engagements are priced
| Model | How it is charged | Fits when | Main risk |
|---|---|---|---|
| Dedicated QA engineer | Monthly rate per engineer | Ongoing product work with a steady release cadence | You still own the test strategy yourself |
| Fixed-scope test cycle | Per project or per test phase | A defined release, migration or audit | Scope disputes at the boundary of what was agreed |
| Managed QA service | Monthly retainer covering the whole function | No QA leadership in-house | You lose visibility into what is not being tested |
| Outcome-based | Tied to an agreed quality measure | You already measure defect escape rate reliably | Poorly designed incentives reward the easiest metric |
Because the deliverable is hard to see, headline rate is the worst available comparison. The useful measure is cost per verified release: the rate multiplied by how much of your team's time goes into triaging false alarms, waiting on flaky tests, and re-testing defects reported as fixed. Two vendors a third cheaper frequently lose on that arithmetic.
On flaky tests specifically: intermittent failures are the fastest way for a testing engagement to become resented. Ask what their policy is for a test that fails one run in ten — quarantine and fix inside a defined window, or rerun until it passes. The second answer quietly teaches your engineers to ignore red builds, which is a more expensive problem than the flakiness itself.
Questions to ask before you sign
- Which of the pricing models above are you proposing for us, and what would make a different one a better fit?
- Can we see a sanitised sample of your test code, and could our own engineers run it?
- Who would be assigned to this engagement, and can we interview them before we commit?
- What is your policy on flaky tests, and how long is the fix window?
- How do you report coverage in a way that is not self-reported?
- What does a defect ticket from your team look like, end to end?
- What is the ramp-up plan for the first two weeks, and what do you need from us to hit it?
- How do you handle a disagreement with our engineering lead about whether something is a bug?
- What is the exit plan, and specifically what will you leave behind in our repository?
- Which engagements have ended, and may we speak to one of those clients?
Red flags
- Coverage percentages quoted before anyone has looked at your product
- A willingness to sign off on a release without an independent quality statement
- Test code that cannot leave their platform
- Any resistance to letting you interview the engineer who will do the work
- Reporting that only ever reassures you
- Reruns used to make intermittent failures disappear from the dashboard
- A rate far below the market, with no explanation of what it excludes
Run a small, real pilot first
The cheapest way to evaluate a testing partner is a paid pilot on real work: one engineer, or one defined test phase, on a live release, with the explicit goal of measuring how they work rather than how many tests they produce. Agree in writing beforehand what the pilot will deliver — a risk assessment of your current coverage, an initial suite committed to your repository, a defect report in your own tracker, and a plain recommendation of what to do next.
Then judge the artefacts, not the presentation. If the pilot's test code is not something your engineers want to keep, the engagement will not improve at scale.
Frequently asked questions
Why not just have developers write the tests?
Developers should own unit tests; they are closest to the code and the feedback loop is fastest there. The gap sits one level above. End-to-end journeys, integrations, and adversarial cases are exactly what the person who wrote the feature is least likely to imagine. A testing partner earns its place where independent judgement is needed, not where a developer already has full context.
Should we keep some testing in-house?
Yes, almost always. Keep ownership of the test strategy, the definition of done, and the release decision. Bring in a partner for capacity, for specialist depth, or to build a capability you will then run yourself. Fully outsourcing the judgement is how organisations end up unable to answer whether a release is safe without asking the vendor who tested it.
How do we measure whether a testing partner is working?
Choose measures that are hard to game: defects found before release that would otherwise have reached customers, defect escape rate after release, the time each release cycle adds, and the share of the suite your own engineers can run and maintain. Test counts and coverage percentages are the easiest numbers to report and the least informative to read.
What is the difference between functional testing and automation testing?
Functional testing describes what is tested: whether the product behaves correctly for its users. Automation describes how it is tested: whether that verification is executed by a script rather than a person. The two are independent, and mixing them up is how teams end up paying for a very large suite that verifies very little.
Where to start
If manual regression is holding your release cadence, or nobody can currently say what your product's test coverage actually is, that is where the conversation should begin. Neuraforz provides QA and functional testing services with engineers who work inside your process, write tests your team keeps, and will say so when a smaller scope is the better call. For the commercial case, start with why QA is the most undervalued investment in software development, then look at how an outsourced QA team changed the release cadence for a fintech platform. Talk to our team and we will scope the smallest engagement that answers your question.