Software changed. Verification has to change with it.

Software used to be written by humans and tested by humans. Now AI writes it at machine speed — and AI cannot be its own final judge.

The AI wrote your code. Who tested it?

EdgeCase is the independent human judgment layer for AI-built software.

Human-led, AI-leveraged: we deploy AI test agents for execution scale, directed by human judgment.

Old worldHumans write → humans test
New worldAI builds → automation checks → humans judge

10+ years testing enterprise software · Telecom, banking & AI background · Direct, accountable delivery

Illustration of a magnifying lens over threads of code, with a hand-drawn ring circling the defects a human reviewer has spotted
Why we exist

AI changed how software is made. It did not remove the need for judgment.

The old QA model treated people as slower test runners. That job belongs to machines now. The work that remains is harder: defining what correct means, challenging the assumptions, and putting a human name behind the release decision.

AI builds →

AI writes code at machine speed.

Automation checks →

Test suites and AI agents run the checks at scale.

Humans judge

We do the part defined checks cannot: decide what matters, challenge ambiguous outcomes, and sign off with evidence.

Testing the tests

The actual job starts where automation stops.

A check still needs a definition of “correct.”

Automation can run a million checks, but it can't decide what “correct” means. Someone has to define the ground truth — especially for non-deterministic AI behavior.

The verifier has to stand outside the loop.

If the same kind of model writes the code and tests it, both can miss the same thing. Independence is not org-chart trivia; it is structural protection against shared blind spots.

Intent is not independent verification.

Product managers define intent, then often verify their own tickets under pressure to ship. That is useful review, but it is still checking your own homework. We arrive with no attachment to the implementation and no incentive to wave it through.

A release needs an accountable human.

In many enterprise and regulated settings, “our agent tested it” is not enough on its own. High-risk releases need accountable human review. We provide the evidence, audit trail, and named sign-off when that standard applies.


Services

Three disciplines, one accountable owner.

Not generic "QA services" — a deliberate blend built for teams that build with AI. Every engagement has one accountable owner, and every finding comes with evidence and a severity you can act on.

The judgment layer

Test strategy, ground truth & sign-off

We define what correct means before the machines run, challenge the assumptions behind the build, and own the evidence-backed release recommendation.

Scope

  • Risk mapping and plain-language ground-truth criteria
  • Domain-informed scenario and adversarial test design
  • Independent review of requirements and generated tests
  • Usability, accessibility, and user-journey judgment

What you get

  • Defect reports with severity and steps to reproduce
  • Evidence pack: screenshots, recordings, data traces
  • A clear human release recommendation — go, or go with fixes
Speed without the flaky suites

Test automation

Regression coverage that runs in your pipeline and stays trustworthy — built, integrated, and handed over so your team can own it.

Scope

  • Web and API automation (Playwright, Cypress, Selenium)
  • CI-ready regression suites with honest coverage reports
  • Flaky-test triage and repair of existing suites
  • Automation health reviews for teams mid-build

What you get

  • Maintainable suites committed to your repo
  • Pipeline integration with clear pass/fail signals
  • Run guides and documentation your team can run with
Execution at machine scale

Human-directed AI test agents

We deploy AI test agents to explore more paths and run more checks — then apply human judgment to their coverage, findings, and blind spots. The agents scale execution; they do not make the release decision.

Scope

  • Validation of AI-generated tests and code output
  • Behavioral testing of LLM and agent-powered features
  • Prompt-aware negative and adversarial testing
  • Human verification loops around automated test output

What you get

  • Risk report on AI behaviors and failure modes
  • Guardrail and fallback recommendations
  • Regression coverage that grows with each release
Why EdgeCase

A decade of testing where failure isn't an option.

We bring 10+ years of QA experience across telecom, banking, pension, crypto, and AI products — software that can't afford to be wrong. We bring that judgment to your release: the rigor to check what matters, and the honesty to say no to a release that isn't ready.

Telecom Banking & Finance Crypto & Digital Assets Healthcare AI & LLM Products
Flagship · AI & LLM products

We've tested live LLM features — not just demos.

Our team's current work includes OpenAI/ChatGPT-based LLM testing for a computer vision AI initiative: hands-on validation of model behavior inside a real, shipping product. That's exactly the discipline we bring to AI teams — behavioral testing of LLM-powered features, prompt-aware negative and adversarial cases, and human verification loops around model output. We don't just sell AI-assisted QA; we practice it.

Sector 01 · Telecom

Customer apps at national scale

Context
Customer-facing web and mobile applications for Canada's largest telecom brands (Rogers, Fido) — products used by millions of subscribers, tested through enterprise release cycles.
Our work
Functional and regression testing across release cycles, internationalization and localization testing (English/French), and WCAG web accessibility testing.
Why it matters for you
Trained judgment for high-volume consumer releases, where a missed defect lands at scale — and where bilingual, accessible experiences are table stakes, not extras.
Sector 02 · Banking & Finance

Where precision is a requirement

Context
Banking platforms (RBC) and pension systems for the Ontario Teachers' Pension Plan — data-heavy, multi-step financial software with compliance-grade expectations, delivered through enterprise consultancy engagements.
Our work
Testing complex financial transactions and workflows, user-acceptance support with business stakeholders, and documentation built to survive audits.
Why it matters for you
Rigor and thoroughness for fintech and regulated products — plus the discipline to document evidence as we go, so your release decision is defensible.
Sector 03 · Crypto & Digital Assets

A secured exchange, tested end to end

Context
Test planning and execution for a highly secured crypto exchange — an environment where authentication and transaction defects cost real money.
Our work
Login and two-factor authentication flows, transaction testing, and gas-fee calculation validation across the exchange's critical paths.
Why it matters for you
Security-critical, money-moving flows tested with a compliance-conscious eye — the same care we bring to any product where trust is the feature.
Automation & tooling

Depth behind the discipline

Selenium Playwright Robot Framework Postman SoapUI SQL

Acceptance automation built with Robot Framework (BDD) that cut testing effort by 50% — alongside API testing with Postman and SoapUI, and ETL/database validation with SQL.

Coverage across web, iOS and Android mobile, REST/SOAP APIs, ETL pipelines, WCAG accessibility, and cross-browser matrices — in Agile and Waterfall teams alike.

The playbook you get — on day one

Test strategy in plain languageWhat we test, what we skip, and why — written, not assumed.
Structured regression suitesCoverage that grows with every release instead of rotting.
A bug-triage rhythmDefects ranked by real severity, not by who shouted loudest.
Release-readiness reportsEvidence-backed go / no-go calls leadership can sign off on.
Automation you actually ownTests in your repo, docs your team understands, no lock-in.
Honest no-go callsThe job is telling you the truth about your release — kindly, firmly.
How it works

From first call to signed-off release.

A steady, transparent rhythm — you always know what's being tested, what's been found, and what happens next. No black box.

Abstract illustration of ascending steps representing a step-by-step QA process
Intro call

Free, 30 minutes. We scope the product, the risks, and what "done" looks like.

Test plan

A written strategy: scope, approach, environments, and timeline — approved by you.

Execution

Human, automated, and AI-assisted testing running side by side.

Reporting

Every defect documented with evidence, severity, and priority.

Re-verify and sign off

Fixes retested, and you get a clear go / no-go — with reasons.


Ways to work together

Start small. Expand on evidence.

Three engagement shapes, one principle: every engagement starts with a written, fixed-scope proposal, so you know exactly what you're getting and what it costs.

Monthly · ongoing

Fractional QA Partner

A QA team without hiring one — steady testing capacity for teams that ship continuously.

  • Dedicated monthly testing capacity
  • Regression suite ownership and maintenance
  • Standing bug-triage rhythm
  • Release sign-offs, every cycle
Ask about a retainer
One-off · pre-release

Launch Readiness

A rigorous pre-release test cycle before the moments that matter — launches, migrations, big refactors.

  • Risk-ranked test plan for the release
  • Full execution pass across devices and data
  • Evidence-backed go / no-go report
  • Post-launch support window
Scope my launch

Early clients typically start with a pilot — it's the fastest way to see the quality of the work, and it de-risks committing to more.

EdgeCase

Human-touch QA consultancy

The team

Experienced testers. Direct access. No hand-offs.

EdgeCase is a human-touch QA consultancy. Our team brings a decade of software quality assurance from telecom and banking enterprises — customer-facing applications used by millions, money-moving financial workflows, a secured crypto exchange, and production LLM features. Manual testing, test automation, release ownership: we've lived all of it, at scales where sloppy isn't survivable.

We founded EdgeCase on a conviction the industry is just catching up to: as AI writes more of our code, thoughtful human verification matters more, not less. Code gets generated fast; trust gets built slowly — release by release, by someone checking the work.

Every engagement has a clear, accountable owner. No unnecessary hand-offs, no black box. We work with teams across North America — remote-first, with close collaboration when it matters.

Questions

Asked often, answered plainly.

Can't AI just test everything?
AI can run an enormous number of checks, and we use it to do exactly that. Defined checks and model-based evaluation can cover many outcomes, but ambiguous and non-deterministic behavior still needs someone to define the ground truth. We direct the agents, test their tests, and make the final call where human judgment is needed.
Isn't this my PM's job?
PM review is important, but a PM defines the intent and is often under pressure to ship it. Asking that same person to independently verify the result is checking your own homework. We stand outside the build loop, challenge the assumptions without attachment, and produce an evidence trail that can support enterprise or regulated review.
Why human judgment in the age of AI?
AI tools accelerate generation and execution, not accountability. Auto-generated tests still check against assumptions — and somebody has to check the assumptions. Humans with domain knowledge spot the contradictory requirement, the confusing flow, and the off-looking data that scripts silently pass.
Will you replace our engineers' testing?
No. Engineers test what they intended; we test what users will experience. We slot in beside your developers, challenge the build with fresh eyes, and feed findings back in the tools your team already uses.
How fast can we start?
Pilots typically begin within two weeks of the intro call — the call, a written test plan and proposal, then testing starts. Retainers and launch-readiness engagements follow the same rhythm.
Remote or on-site?
Remote-first across North America, which keeps costs down and turnaround quick. We stay closely involved through kickoffs, launch weeks, and audits.
Can you work white-label through agencies or recruiters?
Yes. Discrete QA delivery under your agency's brand, straight into your client's workflow, with NDA-friendly reporting. If you're a recruiter or agency with a client that needs QA capacity, this is exactly the partnership model the practice was built for.
What do you need from us to start?
Access to the product (ideally a staging environment), enough context to test like a user — your docs, demos, or a walkthrough — and one decision-maker who can confirm priorities. The test plan covers everything else.
Contact

Start with a conversation.

A focused pilot is the best first step: two weeks, fixed scope, real findings, and a clear recommendation on what comes next.

  • Share the product and the riskTell us what you are shipping, where quality feels uncertain, and when it needs to be ready.

  • Get a written scopeWe define what to test, what to deliver, the timeline, and the fixed pilot cost.

  • Decide from evidenceThe pilot ends with documented findings and a practical release recommendation.

Tell us about your project.

We’ll reply to the email address you provide. Prefer email? [email protected]