API Performance Overhaul
High-traffic API endpoints were degrading under load — p99 DB latency above 820ms, visible user drop-off at key conversion flows. I did a full diagnostic before touching any code: profiled query execution plans, mapped all N+1 patterns, audited cache TTL strategy, and identified three synchronous downstream calls that belonged off the critical path. The fix was systematic: rewrote the data-access layer with explicit batch-fetch queries, added composite covering indexes on the hot query patterns, replaced the self-defeating fine-grained cache TTLs with domain-aligned invalidation, and moved downstream enrichment to async CompletableFuture chains returning after the core response. Outcome: p99 DB latency dropped 40%, API latency dropped 35%, user conversion at the affected flow improved 12%.
$ Architecture
- →JPA query audit: replaced ORM-generated N+1 patterns with explicit JPQL batch-fetch queries — DB round trips on the critical path reduced 6x
- →Composite index design on MySQL: EXPLAIN-driven analysis of the top-10 hot query patterns; covering indexes added to eliminate full-table scans at load
- →Redis caching redesign: domain-aligned TTLs, tag-based invalidation on write mutations, probabilistic early expiration to prevent cache stampede under burst traffic
- →Async enrichment: downstream service calls moved off the critical path using CompletableFuture; core response returned immediately, enrichments applied without blocking the caller
- →Prometheus latency histograms (p50/p95/p99) per endpoint added as deployment-gate signals in CI — regressions caught before production