CASE STUDY

5x Performance Improvement for a Leading Financial Products and Solutions Company

Ashnik Team

  • Industry :
    Banking Infrastructure, Payments
  • Technology :
    Elastic Stack (Elasticsearch, Logstash, Kibana)
  • Engagement :
    Cluster Redesign, Performance Engineering, Managed Services
THE CUSTOMER

A leading financial products and solutions company hosting core banking and payment gateway switch infrastructure for 50+ banks.

THE SOLUTION

Ashnik merged both clusters into a single active-active Elastic deployment, replacing the VIP failover setup with a load balancer across all nodes.

THE CHALLENGE

Two separate active/passive clusters caused downtime on failover, left infrastructure underutilized, and couldn’t absorb 4x expected ingestion growth, without new technology or new budget.

THE RESULT

Indexing rate up 5x, search rate up 5x, and log retention extended from 5 to 8 days, on the same Elastic Stack.

50,000/sec

Indexing rate, up from 10,000/sec

5,000/sec

Search rate, up from 1,000/sec

3,000

Visualizations, up from 1,500

8 days

Log retention, up from 5 days

Customer Overview

The customer is a leading financial products and solutions company managing infrastructure for over 50 banks’ hosted core banking and payment gateway switch operations, spanning 5 Core Banking tenants and 50+ Payment Gateway integrations.

Its observability needs covered UPI transaction monitoring, OTP validation, core banking transactions, payment gateway monitoring, app and infrastructure monitoring, availability and latency tracking, and success/decline transaction analysis.

As a solution provider bound by strict SLAs to its own banking customers, the company needed an observability platform that didn’t just monitor, but let it take preventive action before issues reached end customers.

The Challenge

Downtime on failover

The active/passive setup caused 2 to 3 minutes of downtime whenever failover kicked in.

Underutilized infrastructure

Running active/passive meant a large share of infrastructure sat idle at any given time.

Performance bottlenecks

Log ingestion and search response were delayed during high workload periods.

Need for high availability

The setup had to move beyond active/passive to something genuinely highly available.

4x ingestion growth expected

Ingestion rate was projected to grow more than 4 times across both existing clusters.

No new technology, tight budget

The solution had to work within the existing Elastic Stack investment and a constrained budget.

The Approach

01
Merge two clusters into one
The two separate 3-node Elasticsearch clusters (ES1-ES3 and ES4-ES6) were merged into a single unified 6-node cluster.
02
Active-active, not active-passive
All 4 Logstash + Kibana nodes now run active-active, replacing the previous 2 active / 2 passive split across two separate clusters.
03
Load balancer instead of VIP
A load balancer now distributes traffic across all active nodes, replacing the VIP-based failover in the old design.
04
Same stack, redesigned topology
No new technology was introduced. The performance gain came from re-architecting how the existing Elastic Stack was deployed.
05
Extended retention
Log retention was extended from 5 days to 8 days as part of the redesign.
06
Design, deploy, manage
Ashnik’s role covered designing and deploying the merged cluster, performance enhancement, and ongoing managed services.

Architecture

Before — Two Separate Active/Passive Clusters

cluster1and2 scaled

After — Merged Active-Active Cluster

metrics case study scaled

Scale

METRIC BEFORE AFTER
Cluster topology 2 separate clusters, active/passive 1 merged cluster, active-active
Indexing rate ≈10,000 events/sec ≈50,000 events/sec
Search rate ≈1,000/sec ≈5,000/sec
Visualizations ≈1,500 ≈3,000
Watcher alerts ≈700 ≈1,000
Log retention 5 days 8 days

Outcome

Conclusion

The 5x performance gain didn’t come from new technology or a bigger budget. It came from re-architecting how the existing Elastic Stack was deployed, merging two underutilized active/passive clusters into one fully active cluster with a load balancer in place of VIP failover.

For infrastructure providers bound by strict SLAs, the lesson is architectural: the constraint of “no new technology, tight budget” doesn’t rule out a 5x performance jump. It just means the answer has to come from topology, not spend.