All posts
Load Testing

How to Read API Load Test Results

Read latency, throughput, failures, and server signals together to turn a test report into a useful diagnosis.

API Test Lab4 min read

Part 3 of the API performance testing series. Read latency, throughput, failures, and server signals together to turn a test report into a useful diagnosis.

Find the useful signals: Latency, Errors, Throughput
Find the useful signals: Latency, Errors, Throughput

Begin with completed work

Before looking at response times, confirm the test performed the intended journey. Inspect achieved request rate, successful transactions, and failed requests. A fast response that skips the database or returns an error does not prove your application handled the workload.

Separate attempted requests from successful business operations. For a multi-step checkout, three successful HTTP requests might represent only one completed checkout.

Read latency as a distribution

The median describes the middle request; p95 and p99 describe the slower tail. Compare them together. A stable median and rising p99 may indicate a subset of requests waiting on a slow dependency or queue.

Do not average percentile values from separate workers and call the result a global percentile. Calculate across the combined observations or a compatible histogram. Use the existing latency percentile guide for the basics.

Group failures before investigating

Failure groupUseful next check
Authorization failuresTest account permissions and token lifetime
Rate limit responsesTraffic policy and identity distribution
Server errorsApplication logs and dependency failures
TimeoutsQueues, network path, and upstream duration
Failed body assertionsResponse contract and business outcome

Keep these groups separate. Authentication mistakes in test setup can dominate the report and hide the actual performance behavior. Review a small sample of failed responses without exposing credentials.

Find where throughput stops growing

Compare successive traffic steps. If attempted requests increase while successful throughput stays flat, identify what else changes at that moment. Database connection wait, worker saturation, CPU, or a growing dependency queue can help locate the constraint.

Correlation is a starting point, not proof. High CPU may be the consequence of retry storms rather than the original bottleneck. Use logs or traces to test a specific explanation.

Check the load generator too

The generator must have enough resources to deliver the workload. Inspect its achieved rate, CPU, connections, and network limits. Distinguish client-side waiting from server response time when the tool exposes those timings.

Include warm-up, steady traffic, and recovery as separate intervals. An overall average can hide a brief but important failure during a traffic transition.

Write a decision, not just a report

Summarize the workload, measurement window, pass criteria, actual outcomes, and likely constraint. For example: at the planned rate, successful throughput met demand, but p95 exceeded the agreed target while database connection wait increased.

Choose one follow-up change, repeat the same workload, and compare. For a release workflow, continue to the final article in this series.

Continue the series

1. How to Build an API Load Testing Plan

2. Load, Stress, Spike, and Soak Testing Explained

3. How to Read API Load Test Results

4. API Performance Regression Checklist for Releases

Show all blogs

Share

Start testing your APIs

Try API Test Lab free. No credit card required.

Start free

More from the blog

Read 3 related articles from our latest posts.