Release regression testing for SAP · Salesforce · Dynamics 365 · NetSuite · Veeva

Your vendor picks the release date.
You still have to prove nothing broke.

Lastest runs Playwright regression suites inside your own network against SAP, Salesforce, Dynamics 365, NetSuite and Veeva — capturing visual, structural and perceptual diffs of every screen, as evidence you own.

Runs inside your networkPlaywright tests in your git repoDeterministic replay
Prove what changed on screen · nothing leaves your perimeter

Built for releases you don't control

SAP S/4HANA & ECCSalesforce preview sandboxDynamics 365 F&O & CENetSuite Release PreviewVeeva Vault & Vault CRMPlaywright in your repoSelf-hosted on your infrastructure

The release is dated. The scope is yours. The evidence is on you.

SAP ships a support pack. Salesforce upgrades the preview instances before the release weekend. Microsoft pushes a service update into your UAT sandbox seven days before production. Oracle tells you your NetSuite phase three weeks out. Veeva puts a general release on the calendar three times a year. In every case the job is the same, and Oracle states it plainly in its own documentation: you are not expected to test the new features — you are expected to prove your own configuration still works.

01

Round up the key users and work the spreadsheet.

A test plan in Excel, a room of process owners doing this on top of their day job, screenshots pasted into a document for sign-off, and a defect list chased over email. It works — right up to the release where somebody shortcuts the return-goods scenario because there are only so many hours before the cutover.

  • Not repeatable — the next release starts the whole cycle again
  • Only the happy flow gets tested when the window is short
  • Evidence is assembled by hand, then chased for review
  • Nobody can say later which screen changed, or when
02

Record each process once. Replay it every release.

A key user clicks through order-to-cash, invoice cancellation, period-end close — once. Lastest writes deterministic Playwright code into your repository and captures a baseline of every screen. When the vendor date arrives, the whole pack replays overnight against the sandbox and hands back a screen-by-screen diff.

  • Runs on your infrastructure — nothing leaves your network
  • The same run every time — no model deciding what to click
  • Screen-level diffs catch the renamed tab and the moved field
  • Evidence is generated, not assembled — timestamped per screen
Every other vendor asks you to move your regression suite into their cloud and their format. Lastest runs Playwright inside your network, writes the tests into your repo, and hands you the screenshots.

Not a platform you move into. A runner you keep.

01

It runs inside your network.

Lastest is self-hosted. The runner and the browsers execute on your own infrastructure, next to the system under test, and the screenshots land on storage you control. Production ERP and CRM screens carry customer, patient, employee and financial data — with a self-hosted runner none of it has to cross your perimeter to be regression tested.

  • No callback to our infrastructure at run time
  • Works on isolated network segments with egress blocked
  • Survives a security review because there is nothing to send
02

Your tests are Playwright, in your git repo.

Not a proprietary model repository, not a flow diagram in someone else's SaaS. A readable Playwright file per business process, reviewed in a merge request like any other change, versioned alongside the rest of your landscape. Anyone who has migrated a regression suite off a closed platform knows why this line matters.

  • Read them, fork them, diff them, keep them if you leave
  • Runs anywhere Node.js runs — no Lastest required to execute
  • Selector strategy is visible in the file, not hidden in a product
03

It diffs the actual screen — visually, structurally, perceptually.

A passing assertion tells you a field had a value. It does not tell you the tab was renamed, the field moved to another card, the Fiori tile stopped rendering, or the dispatcher console quietly became something else. Lastest compares the rendered screen three ways and shows you the region that moved.

  • Pixel comparison for exact change detection
  • Structural (SSIM) comparison for layout shifts
  • Perceptual comparison for what a human eye would notice
  • Text-region-aware diffing so dynamic document numbers don't fire
04

AI records once. Replay is deterministic.

A model is involved exactly once: when the process is recorded or a step is repaired, it writes the Playwright code and the selector strategy. At replay time there is no model in the execution path — the committed code executes, and the same run against the same build produces the same evidence. If it fails, it failed for a reason you can read in a file, not because a runtime agent made a different choice than last month.

  • Authoring time: AI writes the test
  • Replay time: zero inference, zero variability
  • Reproducible runs are what makes the output evidence
05

The screenshots are your sign-off evidence.

Every run produces a timestamped before/after image per screen, tied to the process, the environment and the build — generated automatically rather than exported by hand into a document. It is the artefact the spreadsheet method has always been trying to produce, without the afternoon of manual PDF exports, and it is exportable for change control, UAT sign-off and audit.

  • Machine-generated, timestamped, traceable to the run
  • Replaces the manual export-and-compare step entirely
  • Stored in your environment, attachable to your change record

What you actually get back. Screen by screen.

Nobody in this category shows you the output. Here is ours.

01 · Record

A key user clicks through the process. That's the test.

Open the recorder, point it at the sandbox, and walk the business process the way a process owner walks it: create the order, release the delivery, cancel the invoice, close the period. Stop recording. Lastest emits deterministic Playwright code with a layered selector strategy, and a baseline screenshot for every screen you passed through.

  • No access to the system's source or metadata required — it tests the running application
  • Works against SAP GUI in a browser, Fiori, Lightning, F&O, NetSuite and Vault UIs
  • The generated file lands in your repository, ready for review
Recorder timeline showing each captured step of a recorded business process

02 · Diff

After the release, every screen comes back scored.

The pack replays against the upgraded environment and each screen is compared to its baseline with three diff engines — pixel, structural and perceptual. Unchanged screens pass silently. Changed screens come back with the region highlighted and a before/after slider, so a process owner can rule in five seconds whether the change is expected or a defect.

  • Every diff is reviewable and dispositioned — accepted as the new baseline, or raised
  • Run history shows which release introduced which change
  • Export the run as the evidence pack for your change record
Diff slider comparing the baseline screen against the screen captured after the vendor release

Impact analysis tells you what to test. Lastest tells you what changed.

If you run Tricentis LiveCompare, Panaya or Worksoft Impact, keep them. They analyse metadata, code and usage to narrow your regression scope, and they are good at it — Lastest does not do that analysis and does not claim to. It is the layer after the scope is set: it executes the processes and compares the screens, so the answer to “did anything actually change for the user?” is an image, not an inference. The two halves fit together; owning one has never told anyone what the screen looked like afterwards.

Every one of them sets your date. Pick yours.

The cadence, the vocabulary and the window are different in each. The method is the same.

Feature packs, support packs and system conversion milestones — with screen-level diffs of both SAP GUI and Fiori, inside your own landscape.

SAP S/4HANA regression testing →

Baseline Lightning pages, Flows and page layouts in the preview sandbox, then diff them against production after the release weekend.

Salesforce regression testing →

Your sandbox takes the service update seven days before production. Replay your recorded F&O and CE business processes in that window.

Dynamics 365 regression testing →

Retire the Release Preview test plan spreadsheet. Diff saved searches, custom forms and SuiteFlow screens between preview and your live account.

NetSuite regression testing →

Veeva performs IQ and OQ. You still owe PQ against your own configuration — replayed on the prerelease Vault, evidence captured automatically.

Veeva Vault regression testing →

Questions your security, QA and validation reviewers will ask

Can Lastest run entirely inside our own network?
Yes — that is the default deployment. Lastest is self-hosted: you run the containers on your own infrastructure, inside your perimeter, and point them at your sandbox, preview or prerelease environment. There is no callback to our infrastructure at run time, so it works on an isolated network segment and on hosts with no outbound internet access. Nothing about a production ERP or CRM screen has to leave your firewall for the regression run to happen.
Where do the screenshots and test data live?
On storage you control. Baselines, run screenshots, diff images and logs are written to the volume or object store you configure in your own environment. Because production ERP and CRM screens routinely contain customer, patient, employee and financial data, this matters for data residency, GDPR and GxP scope: the evidence never enters a vendor cloud unless you deliberately put it there. If your security review needs it, the runner can be operated with egress blocked entirely.
What happens to our regression suite if we stop paying?
You keep it. Your tests are plain Playwright code committed to your own git repository, and your baselines and screenshots are files in your own storage. There is no proprietary model repository to be locked out of and no export process to negotiate. If you leave, the suite still runs — it is Playwright, so any team with Node.js can execute it. That is the opposite of the usual arrangement in this category, where the tests live in the vendor format and leaving means starting the suite again from zero.
Does AI run when the tests replay?
No. AI is used at authoring time only — when you record a business process or repair a step, a model watches the interaction and writes the Playwright code and selector strategy. After that the model is out of the picture. Replay executes the committed code with no inference in the execution path, so the same run against the same build produces the same result and the same evidence. This is deliberate: an auditor, a Basis lead and a validation lead all need a run to be reproducible, and a model deciding at run time what to click is not reproducible.
How is this different from change-impact analysis tools?
It is complementary, not competing. Change-impact tools such as Tricentis LiveCompare and Panaya analyse metadata, code and usage to tell you which objects a release touches, so you can scope your regression. Lastest does not do that analysis and does not claim to. Lastest is what runs after the scope is set: it executes the business processes in a real browser and compares the rendered screens, so you can show what actually changed on screen rather than only which objects were flagged. Teams that own an impact-analysis tool typically use it to choose the scope and use Lastest to prove the outcome.
Do we need a git repository, and do our testers need to write code?
A git repository is strongly recommended because it is what makes the suite yours and gives you review history on every test change, but the runner can also read tests from a mounted directory if your organisation has no repo for QA assets. Test authors do not write code: a functional consultant or key user records the process by clicking through it, and the Playwright file is generated. The code is there so that engineers can read it, review it in a merge request and edit it when a process changes.
Why compare screens instead of asserting on fields and APIs?
Because a large share of release breakage is not an error. A tab gets renamed, a field moves to another card, a layout collapses on a Fiori tile, a button stops rendering for one role — assertion-based and API-level tests pass through all of that untouched, and the first person to notice is a user in production. Lastest compares the actual screen three ways: pixel comparison, structural (SSIM) comparison for layout shifts, and perceptual comparison for changes a human eye would register. Text-region-aware diffing keeps dynamic data like document numbers and timestamps from firing false positives.
What is a release-readiness review?
A 30-minute working session, not a product pitch. We map your next vendor-set release date — the support pack, the release weekend, the service update, the Release Preview window, the general release — to a concrete regression scope: which business processes have to be proven, which screens carry the evidence, and what the run looks like on your infrastructure in the window you actually have. You leave with the scope written down whether or not you go further with us.

You know the date.
Let's scope the regression.

Thirty minutes. We map your next vendor-set release — support pack, release weekend, service update, Release Preview, general release — to a concrete regression scope: the processes that have to be proven, the screens that carry the evidence, and what the run looks like on your infrastructure inside the window you were given. You keep the scope either way.

Book a release-readiness review →
release-run

# runs on your host, against your sandbox

$ lastest run --suite regression/order-to-cash

 

42 processes · 318 screens · baseline 2026-07-14

311 screens unchanged

! 7 screens changed · review required

 

# evidence written to your storage

$ ls artifacts/2026-07-29/

before/ after/ diff/ run.json