For some background, see The Past Story about using UI-Tars for AI Testing, which shows practical AI applications in test automation using UI-Tars and Midscene.js.
What you’ll get from this blog: In just 5 minutes, the time for a quick smoke or a cup of coffee, I’ll show you a practical example of handing things over to an E2E-Test AI agent and watching it autonomously carry out full end-to-end testing jobs.
Table of Contents
- Let's see the Demo first
- Introduction
- What can the E2E Test AI Agent do as my new team member?
- E2E Test AI Agent Architecture
- New Format of the E2E Test Case
- Generated test cases are robust, much less flaky than human-written
- How This Generated Code Differs from Others
- Why I use AI? Practical Applications in E2E Testing
- I’d love to hear from you!
Let's see the Demo first
00:00 Interactive with AI to generate test
01:50 Rerun the generated test from the beginning
Key Features of this Demo E2E-Test AI Agent
Interactive Acceptance Criteria Input: Enter your acceptance criteria step by step. Every action is visualized and fully reversible, so you can adjust at any point.
Image-Based Interaction: Use images as input for interactions, for example, clicking on some Canvas element on the screen.
Test Data Configuration: Beyond basic browser operations, the AI agent allows you to configure test data, such as setting the current session to a registered user.
Test Generation & Compatible with existing E2E test framework: All actions are automatically converted into Playwright-based tests with your in-house E2E Test framework. You can export them and put them as part of your existing E2E automation test suite.
Smart, Self-healing, Reliable Execution: Generated tests automatically handle waits for navigation, waits for necessary HTTP/GraphQL requests, and other timing issues to ensure tests are not flaky.
Automated Feature & Scenario Generation: generated codes are automatically structured into high-level Feature and Scenario definitions.
Introduction
In software engineering, the User Acceptance testing via the E2E test approach is the last line of defense in Functional Quality before delivering to customers. It’s the closest to real user behavior and business, and even though it’s costly to write and maintain, it’s the most trustworthy way to ensure what reaches customers actually works.
In practice, even a highly experienced SDET/QA engineer can spend over 30 minutes just to write and automate a single test case, not to mention the ongoing maintenance cost. When a new feature rolls out, it usually starts with the PM laying out the requirements, the UI designer shaping how it looks, and then the engineers building it. In teams with decentralized QA, engineers often end up as the last stop and are encouraged to turn the acceptance criteria into test cases and manually execute them or automate them with some tools like Playwright.
The problem usually arises at the final stage: It’s perfectly fine for engineers to own end-to-end quality, that’s how good products ship. The pain comes from the fact that with most E2E automation frameworks/approaches today, creating and maintaining those tests is still a huge time sink: It isn’t only writing tests, it’s tracking down the right element locators and then babysitting those tests every time the UI shifts...
But look, the great collaboration & the (roughly) clear acceptance criteria ran through the entire workflow! What if we could bring PMs, UI designers, and anyone into the process, and slightly adjust the format of acceptance criteria so they’re directly executable as automated E2E tests that engineers can run?
What can the E2E Test AI Agent do as my new team member?
My vision is simple: Please help to strip away the busywork as much as possible.
- If E2E tests can be generated, they must originate from business requirements, not from any piece of your product code.
Human: interact with an AI agent to state the acceptance criteria and assertions in the way you’d talk to any colleague.
E2E-Test AI Agent: interact back with the human to execute those criteria and verify the assertions.
As has been said, E2E-Test AI Agents are your new colleagues, not just tools.
Let’s explore the E2E-Test AI Agent I built at Creative Fabrica, a digital creative platform in Amsterdam offering fonts, graphics, craft resources, and AI design tools, where I currently serve as Head of QA.
E2E Test AI Agent Architecture
Maybe this is the part my readers are really interested in!
Key Component/Layers
- Layer 1: From Human Language to Defined Tools
Translate natural-language acceptance criteria into structured actions and assertions that can be executed, and get different format outputs.
- Layer 2: Framework Integration
Wrap your existing E2E test framework, Midscene.js functions, and Playwright native functions as LLM tools, and keep all LLM Tools can share a single Playwright Browser Context.
- Layer 3: AI-Driven Planning
Midscene.js orchestrates the rest: planning AI steps, interpreting the current screenshot and HTML DOM, and deciding the next best action autonomously.
New Format of the E2E Test Case
In the current paradigm, a test case created and autonomously executed by an E2E-Test AI Agent contains 3 core metadata layers:
Acceptance Criteria (Human Input):
The intent, expressed directly in natural language by humans.Executable Test Code (Auto-generated):
Playwright-based code that integrates seamlessly with your existing E2E framework.Element Locator Cache (AI-generated by Midscene.js):
Cached mappings of HTML elements for fast execution. Only when the cache expires or is missing does the agent call the LLM again to resolve new locators.
With these three, a test case is no longer a static artifact, but a self-adaptive entity that balances human intent, system execution, and AI-assisted resilience.
Here are the exported files from the demo in the video:

Generated test cases are robust, much less flaky than human-written
While I don’t have solid numbers to share yet, I can show you something better, the actual generated code from the demo:)
PS: the Describe title and test title are also generated by LLM :) I know I know... Finding a proper and standard-compliant name every day for a human is not that easy :))
import {
expect
} from '@playwright/test'
import {
tags,
cfAITest as test
} from '@cf/pw-baseline'
test.describe('Download a freebie product', () =>
test(
'As a registered user, I can download a freebie product and confirm its successful download',
tags({
owner: 'ai',
criticality: 'low',
id: 'ai',
}),
async ({
cleanPage,
testUser,
waitGraphQL,
aiTap,
aiAssert,
aiWaitFor,
}) => {
await Promise.all([
cleanPage.waitForURL(url => {
return url.pathname.endsWith('/')
}),
waitGraphQL('/query', 'query siteData'),
waitGraphQL('/query', 'query cfMenuUser'),
waitGraphQL('/query', 'query user'),
waitGraphQL('/query', 'query userAuth'),
cleanPage.goto("/")
])
await testUser("REGISTERED", {
attachToPage: true,
})
await aiWaitFor("Freebies from the top menu is visible")
await Promise.all([
cleanPage.waitForURL(url => {
return url.href.includes("/freebies/")
}),
aiTap("Freebies from the top menu")
])
await aiWaitFor("1st product is visible")
await Promise.all([
cleanPage.waitForURL(url => {
return url.href.includes("/product/autopub-graphic/")
}),
aiTap("1st product")
])
await aiTap({
prompt: "this image",
images: [{
name: 'this image',
url: 'data:image/png;base64,..<I removed this too big base64 string>....'
}]
})
await aiWaitFor("the success download popup is visible")
await aiAssert("the success download popup is visible")
},
))
How This Generated Code Differs from Others
Based on the suggestions from Midscene.js, the code generation uses Instant Actions frequently (e.g.,
aiTap,aiInput) instead ofaiAction. This approach has indeed made the LLM actions and planning more stable than before.The code generation can leverage features from our existing E2E test framework, such as generating test users and attaching user cookies to the current browser session.
The code generation can leverage native Playwright features and apply Playwright Best Practices autonomously, such as
page.waitForURL()and usage ofPromise.all(), etc.
Why I use AI? Practical Applications in E2E Testing
I'm using AI is not about following trends, but it’s about solving real-world problems over the past decade. Based on the demo video of the E2E Test AI Agent I showcased, I see the following practical applications:
Shifting Testing Earlier in the Design Phase
- Traditionally, QA engineers would design test cases based on acceptance criteria provided by PMs (eg: in TestRail).
- With this AI Agent, PMs and UI designers can directly write acceptance criteria against the E2E-Test AI Agent while designing features, or the tool can even pull acceptance criteria from Jira.
- These acceptance criteria become executable tests automatically, allowing developers to run tests as they implement features and update criteria dynamically.
Optimizing Developer/QA Time
- Typically, Developers/QA often handle end-to-end testing at the final stage.
- With this AI Agent, they no longer need to write test code or maintain it when locators change.
- Developers simply interactively input acceptance criteria, almost like doing pair testing with a manual QA engineer, but the AI handles the test execution autonomously.
I’d love to hear from you!
Feel free to like, comment, or share this blog, your feedback means a lot!
@software{Midscene.js,
author = {Zhou, Xiao and Yu, Tao},
title = {Midscene.js: Let AI be your browser operator.},
year = {2025},
publisher = {GitHub},
url = {https://github.com/web-infra-dev/midscene}
}

SOCIAL SHARE CARD GENERATOR