Introduction
A functional test asks whether a feature works. A load test asks whether it still works, fast enough, when many people use it at once. The answer is often different: a checkout that takes 50 ms for one user can take nearly four times as long for a hundred, without any code changing.
This guide uses k6, Grafana's open-source load testing tool. Tests are JavaScript files, and the same script runs on a laptop or in CI. You'll run three kinds of test against a small shop API on your own machine:
- A smoke test: one user, to prove the script and the API work.
- A load test: a steady number of users, with thresholds that decide pass or fail.
- A stress test: users keep increasing until the API gets too slow, then the test stops.
Every command and output below comes from our runs on 2026-09-15, on a laptop with an Intel Core i7-1270P (12 cores, 16 threads) and 32 GB of RAM, running Windows 11, k6 2.0.0, and Node.js 20.20.2. Summaries are trimmed to the lines discussed.
Before you start: only test what you own
A load test is a flood of requests. Aimed at someone else's site, it can slow the site down for real customers, run up their hosting bill, and look exactly like an attack. Run load tests only against systems you own or have written permission to test, and tell whoever runs the infrastructure before you start.
That's why this guide doesn't use a public demo site. The system under test runs on your machine, and nobody else is affected.
The system under test: a tiny shop API
Save this as server.mjs. It needs Node.js and nothing else.
// A tiny shop API for load testing on your own machine. No dependencies: node server.mjs
// GET /products is fast. POST /checkout waits for one of POOL_SIZE simulated database
// connections, each busy for QUERY_MS, so requests queue once load passes the pool's capacity.
import http from "node:http";
const PORT = Number(process.env.PORT ?? 3456);
const POOL_SIZE = Number(process.env.POOL_SIZE ?? 4);
const QUERY_MS = Number(process.env.QUERY_MS ?? 40);
const products = Array.from({ length: 20 }, (_, i) => ({ id: i + 1, name: `Product ${i + 1}`, price: 5 + i }));
let busy = 0;
const waiting = [];
function withConnection(work) {
return new Promise((resolve) => {
const run = () => {
busy++;
setTimeout(() => {
busy--;
resolve(work());
const next = waiting.shift();
if (next) next();
}, QUERY_MS);
};
if (busy < POOL_SIZE) run();
else waiting.push(run);
});
}
let orders = 0;
const server = http.createServer(async (req, res) => {
const send = (status, body) => {
res.writeHead(status, { "content-type": "application/json" });
res.end(JSON.stringify(body));
};
if (req.method === "GET" && req.url === "/products") return send(200, products);
if (req.method === "POST" && req.url === "/checkout") {
let raw = "";
for await (const chunk of req) raw += chunk;
let items;
try {
items = JSON.parse(raw).items;
} catch {
return send(400, { error: "Body must be JSON" });
}
if (!Array.isArray(items) || items.length === 0) return send(400, { error: "items must be a non-empty array" });
const order = await withConnection(() => ({ orderId: ++orders, items: items.length }));
return send(201, order);
}
send(404, { error: "Not found" });
});
server.listen(PORT, () => console.log(`Shop API on http://localhost:${PORT} (pool ${POOL_SIZE}, query ${QUERY_MS} ms)`));The database is simulated: there isn't one. POST /checkout waits for one of four "connections", and each one stays busy for 40 ms. That's the part to remember, because it sets the limit. Four connections that each finish a query every 40 ms can complete at most 4 × 25 = 100 checkouts a second. Real systems have the same kind of limit in a connection pool, a thread pool, or a rate-limited downstream service; they're just harder to see.
Start it in its own terminal:
node server.mjsShop API on http://localhost:3456 (pool 4, query 40 ms)Install k6
k6 is a single program, not an npm package. Install it with your system's package manager, following the installation page, for example winget install k6 --source winget on Windows or brew install k6 on macOS. Then check the version:
k6 versionk6.exe v2.0.0 (commit/8c3be52cc1, go1.26.3, windows/amd64)k6 2.0 removed some older commands and flags, such as k6 login and --no-summary. If a tutorial uses them, check the 2.0 release notes for the replacement.
k6 can send anonymous usage statistics to Grafana. Every command in this guide adds --no-usage-report to turn that off.
Step 1: a smoke test with one user
Before adding load, check that the script works and the API answers correctly. Save this as smoke.js:
import http from "k6/http";
import { check, sleep } from "k6";
const BASE = __ENV.BASE_URL || "http://localhost:3456";
export const options = {
vus: 1,
duration: "10s",
thresholds: {
http_req_failed: ["rate<0.01"],
checks: ["rate==1"],
},
};
export default function () {
const products = http.get(`${BASE}/products`, { tags: { name: "products" } });
check(products, { "products: status 200": (r) => r.status === 200 });
const order = http.post(`${BASE}/checkout`, JSON.stringify({ items: [1, 2] }), {
headers: { "Content-Type": "application/json" },
tags: { name: "checkout" },
});
check(order, {
"checkout: status 201": (r) => r.status === 201,
"checkout: has an order id": (r) => typeof r.json("orderId") === "number",
});
sleep(1);
}A few k6 words:
- A VU (virtual user) runs the
defaultfunction in a loop.vus: 1means one user. - One pass through the function is an iteration.
sleep(1)is the pause a real person takes between actions. Without it, one VU sends requests as fast as the API answers, which isn't what one person does. - A check records whether a condition held. Unlike an assertion in a functional test, a failed check doesn't stop the test.
- A threshold is the pass-or-fail rule for the whole run. Here: fewer than 1% of requests may fail, and every check must pass.
- The
nametag groups requests, so results can be read per endpoint.
Run it:
k6 run --no-usage-report smoke.js █ THRESHOLDS
checks
✓ 'rate==1' rate=100.00%
http_req_failed
✓ 'rate<0.01' rate=0.00%
█ TOTAL RESULTS
checks_total.......: 30 2.85448/s
checks_succeeded...: 100.00% 30 out of 30
checks_failed......: 0.00% 0 out of 30
✓ products: status 200
✓ checkout: status 201
✓ checkout: has an order id
HTTP
http_req_duration..............: avg=24.3ms min=0s med=22.99ms max=56.27ms p(90)=48.61ms p(95)=50.61ms
http_req_failed................: 0.00% 0 out of 20
http_reqs......................: 20 1.902987/sTen iterations, twenty requests, every check passed, and k6 exited with code 0. A smoke test that fails is cheap to fix: a wrong URL, a missing header, or a broken endpoint shows up here, not twenty minutes into a load test.
Step 2: a load test with thresholds
Now the question the business actually asks: does checkout stay fast at the traffic we expect? Say the goal is that 95% of checkouts finish in under 200 ms, and 95% of product listings in under 100 ms. Save this as load.js:
import http from "k6/http";
import { check, sleep } from "k6";
const BASE = __ENV.BASE_URL || "http://localhost:3456";
const TARGET = Number(__ENV.VUS || 20);
export const options = {
stages: [
{ duration: "20s", target: TARGET },
{ duration: "40s", target: TARGET },
{ duration: "10s", target: 0 },
],
thresholds: {
http_req_failed: ["rate<0.01"],
"http_req_duration{name:products}": ["p(95)<100"],
"http_req_duration{name:checkout}": ["p(95)<200"],
},
};
export default function () {
const products = http.get(`${BASE}/products`, { tags: { name: "products" } });
check(products, { "products: status 200": (r) => r.status === 200 });
const order = http.post(`${BASE}/checkout`, JSON.stringify({ items: [1, 2] }), {
headers: { "Content-Type": "application/json" },
tags: { name: "checkout" },
});
check(order, { "checkout: status 201": (r) => r.status === 201 });
sleep(1);
}The stages ramp up to the target number of users over 20 seconds, hold for 40 seconds, and ramp down. http_req_duration{name:checkout} is a threshold on checkout requests only, selected by the tag. The number of users comes from an environment variable, so the same script runs at any level.
Run it with 20 users:
k6 run --no-usage-report -e VUS=20 load.js █ THRESHOLDS
http_req_duration{name:checkout}
✓ 'p(95)<200' p(95)=61.59ms
http_req_duration{name:products}
✓ 'p(95)<100' p(95)=1.08ms
http_req_failed
✓ 'rate<0.01' rate=0.00%
HTTP
http_req_duration..............: avg=24.56ms min=0s med=20.59ms max=94.1ms p(90)=50.54ms p(95)=53.35ms
{ name:checkout }............: avg=48.58ms min=39.44ms med=47.73ms max=94.1ms p(90)=53.36ms p(95)=61.59ms
{ name:products }............: avg=539.7µs min=0s med=541.29µs max=1.75ms p(90)=946.68µs p(95)=1.08ms
http_req_failed................: 0.00% 0 out of 2126
http_reqs......................: 2126 30.259039/sAll three thresholds pass. Look at the first http_req_duration line: its p(95) is 53.35 ms, and it mixes both endpoints. Half the requests are product listings that take about half a millisecond, so they pull every overall figure down. The checkout line underneath is the one that matters for the goal. This is why the thresholds are per endpoint.
Then we ran the same script at 80 and 100 users:
| Users | Requests per second | Checkout median | Checkout p(95) | Iteration p(95) | Thresholds |
|---|---|---|---|---|---|
| 20 | 30.3 | 47.73 ms | 61.59 ms | 1.06 s | Pass |
| 80 | 120.0 | 47.01 ms | 52.22 ms | 1.05 s | Pass |
| 100 | 136.6 | 184.66 ms | 198.04 ms | 1.19 s | Pass |
| 120 | 141.0 | 420.62 ms | 447.1 ms | 1.44 s | Fail |
| 150 | 143.0 | 847.79 ms | 899.48 ms | 1.9 s | Fail |
At 80 users, checkout was as fast as with 20. At 100, the median jumped from 47 ms to 185 ms, and the 95th percentile landed 2 ms under the goal. Nothing failed and every threshold passed, so a CI job would have gone green, with checkout four times slower than a few minutes earlier. A threshold only catches what it's set to catch.
At 120 users the checkout threshold failed, and at 150 checkout took nearly a second. Product listings stayed around a millisecond the whole time: only the endpoint behind the pool slowed down. Still no request failed. A slow system isn't a broken one, which is why a response-time threshold is needed as well as http_req_failed.
When a threshold fails, k6 marks it with ✗, logs an error, and exits with a non-zero code. Both failing runs ended like this:
http_req_duration{name:checkout}
✗ 'p(95)<200' p(95)=447.1ms
time="2026-09-15T22:56:49+03:00" level=error msg="thresholds on metrics 'http_req_duration{name:checkout}' have been crossed"The exit code was 99 for both. CI treats any non-zero exit code as a failure, so the job goes red without any extra configuration.
Reading the summary
A few lines in the k6 summary are easy to misread:
p(95)is the 95th percentile: 95% of requests were at least this fast. It's usually a better goal than the average, because the average hides a slow minority, and they're real users too.http_reqsper second is throughput. At 100 users it rose to 136.6 a second, not 200: two requests per iteration, but each iteration took longer.iteration_durationis one pass through the function, includingsleep(1). At 100 users its 95th percentile grew from 1.05 s to 1.19 s. The users were waiting for checkout.
That last point matters. In this script, each VU waits for a response before sending the next request. When the API slows down, the VUs slow down too, so the load k6 sends drops just when the API is struggling. You can see it in the table: from 120 to 150 users, 25% more users, requests per second rose by only 1.4%. Real visitors don't wait for each other.
k6's documentation calls this a closed model, and names the problem it causes coordinated omission: an overloaded system looks better than it is. The arrival-rate executors, constant-arrival-rate and ramping-arrival-rate, are open models. They start iterations at a set rate, however slow the responses are. Use one when the question is "what happens at 100 checkouts a second?" rather than "what happens with 100 users?".
Step 3: find the limit with a stress test
The load test answered a yes-or-no question for one number of users. A stress test asks where the limit is. Instead of running the load test again and again, keep adding users and stop as soon as the goal is missed. Save this as stress.js:
import http from "k6/http";
import { check, sleep } from "k6";
const BASE = __ENV.BASE_URL || "http://localhost:3456";
// Keep adding users until checkout gets too slow, then stop instead of running the whole ramp.
export const options = {
stages: [{ duration: "3m", target: 300 }],
thresholds: {
"http_req_duration{name:checkout}": [{ threshold: "p(95)<200", abortOnFail: true, delayAbortEval: "10s" }],
},
};
export default function () {
const order = http.post(`${BASE}/checkout`, JSON.stringify({ items: [1, 2] }), {
headers: { "Content-Type": "application/json" },
tags: { name: "checkout" },
});
check(order, { "checkout: status 201": (r) => r.status === 201 });
sleep(1);
}The ramp goes from 0 to 300 users over three minutes. abortOnFail: true stops the test as soon as the threshold fails, and delayAbortEval: "10s" waits ten seconds before the first check, so a few slow requests at the start can't stop it early.
k6 run --no-usage-report stress.js █ THRESHOLDS
http_req_duration{name:checkout}
✗ 'p(95)<200' p(95)=249.22ms
EXECUTION
vus............................: 102 min=2 max=102
running (1m02.0s), 000/300 VUs, 2918 complete and 103 interrupted iterations
default ✗ [ 34% ] 049/300 VUs 1m02.0s/3m00.0s
time="2026-09-15T22:59:45+03:00" level=error msg="thresholds on metrics 'http_req_duration{name:checkout}' were crossed; at least one has abortOnFail enabled, stopping test prematurely"The test stopped after 62 seconds, a third of the way through, with 102 users running, and exited with code 99. That matches the load tests: 100 users passed by 2 ms and 120 failed.
Two things to keep in mind when reading a stress result:
- The threshold is evaluated over the whole run so far. The p(95) of 249.22 ms includes every fast request from the first minute, so it crosses 200 ms a little after checkout really started slowing down. Treat the number of users at the stop as an upper estimate, and confirm it with a load test just below it.
- The limit is where the queue starts, not where requests fail. No request failed here. In a real system, requests usually start timing out or failing some time after response times climb, so waiting for errors finds the limit far too late.
Exercise: find your machine's limit
The API reads two settings from environment variables: POOL_SIZE, the number of simulated database connections (default 4), and QUERY_MS, how long each query takes (default 40). Changing them is like giving the database more connections or a faster query.
-
Before running anything, predict how many users the default API handles before checkout's 95th percentile passes 200 ms. Use only the two settings and the fact that each user sends one checkout and then sleeps for a second.
-
Restart the API with twice as many connections and predict the new limit.
POOL_SIZE=8 node server.mjsIn PowerShell, set the variable first:
$env:POOL_SIZE = 8; node server.mjs. -
Run
stress.jsagainst it and compare the number of users at the stop with your prediction. -
Add a threshold to
load.jsso the run also fails when any check fails. Then prove that this threshold, and nothttp_req_failed, is what fails the run: add a check that the response toGET /productslists 21 products. It lists 20, but the status is still 200.
Your numbers will differ from ours. What should match is the pattern: how the limit moves when you change the pool.
Answer key
1. The default limit. Each connection finishes a query every 40 ms, so it can handle 1000 ÷ 40 = 25 checkouts a second. Four connections handle 4 × 25 = 100 a second. In stress.js, each user sends one checkout and then sleeps for a second, so while checkout is fast each user sends a little under one checkout a second. The prediction is about 100 users. Above that, requests wait in the queue, and the wait grows with every extra user. Our stress run stopped at 102 users.
2. Twice the connections. Eight connections handle 8 × 25 = 200 checkouts a second, so the prediction is about 200 users. Our run with POOL_SIZE=8 stopped after exactly two minutes, at 200 users:
http_req_duration{name:checkout}
✗ 'p(95)<200' p(95)=200ms
running (2m00.0s), 000/300 VUs, 11077 complete and 200 interrupted iterations
time="2026-09-15T23:02:16+03:00" level=error msg="thresholds on metrics 'http_req_duration{name:checkout}' were crossed; at least one has abortOnFail enabled, stopping test prematurely"Doubling the pool doubled the limit. The p(95) shows as 200ms and still fails, because the threshold is "under 200" and the value is rounded for display.
The general rule: capacity is the number of things that can work at once, times how many each can finish per second. Whatever has the smallest capacity, here the connection pool, sets the limit for the whole endpoint. That's why product listings stayed at a millisecond while checkout slowed down.
3. How your numbers can differ. Real queries don't take exactly 40 ms on a busy machine, and k6 shares the CPU with the API, so your stop may come a few users earlier or later. If your limit is far below the prediction, check whether something else on the machine is using the CPU.
4. Failing on checks. Add checks: ["rate==1"] to the thresholds, next to http_req_failed. With a check for 21 products, only the checks threshold failed, and k6 exited with code 99:
checks
✗ 'rate==1' rate=50.00%
http_req_failed
✓ 'rate<0.01' rate=0.00%
✗ products: 21 products
↳ 0% — ✓ 0 / ✗ 5Every request succeeded, so http_req_failed passed. Without the checks threshold, the run would have passed while every product list was wrong.
If you tried a request to a URL that doesn't exist instead, you'll have seen a different result. We sent one request per iteration to /prodcts, a typo, and both thresholds failed:
checks
✗ 'rate==1' rate=50.00%
http_req_failed
✗ 'rate<0.01' rate=50.00%By default, k6 counts every response with a status of 400 or above as a failed request, so a 404 fails http_req_failed whether or not you check it. A check is what catches a wrong answer that arrives with a successful status.
What this doesn't tell you
- Your production numbers. A laptop running both k6 and the API shares its CPU between them. Real systems have networks, real databases, caches, and other traffic. Load test in an environment built like production, with realistic data.
- Anything about the browser. k6's HTTP tests measure the API. Page rendering, JavaScript, and images need other tools, such as k6's browser module or Lighthouse.
- Whether 200 ms is the right goal. Thresholds come from what users and the business need, not from what the system happens to do today.
- Long-running problems. Memory leaks and slowly filling disks need soak tests that run for hours, not a minute.
Conclusion
Start with one user to prove the script works, then add load with a goal written as thresholds, one per endpoint that matters. Read the percentiles, not the averages, and watch the iteration time, because a closed model hides overload. When a threshold fails, k6 exits with a non-zero code, so the same script can guard every release in CI.