CASE STUDY

Log Monitoring at Scale for a Leading Payment Gateway Provider

Ashnik Team

  • Industry :
    Payments
  • Technology :
    Elastic Stack (Elasticsearch, Logstash, Kibana), NGINX+
  • Engagement :
    Platform Design, Managed Services
THE CUSTOMER

One of India’s leading online payment gateway companies, running Payment Aggregator, Biller Network, and Recurring Payment services.

THE SOLUTION

Ashnik designed and deployed a dedicated ELK log-monitoring stack behind an NGINX+ load-balancing layer, taken from UAT to production with 24×7 managed support.

THE CHALLENGE

The customer needed an experienced Elastic Stack partner to build log monitoring across 25+ microservices, with no easy way to trace a transaction or monitor latency and HTTP status per service.

THE RESULT

~500 GB of log data processed daily at ~16,000 events/sec (1.4 billion events/day), with 10-day retention and seamless troubleshooting across microservices.

16,000/sec

Indexing rate

1.4 Billion

Events processed per day

500 GB

Log data ingested daily

10 days

Log retention period

Customer Overview

The customer is one of India’s leading online payment gateway companies, delivering banking and merchant website transactions through a digital network of agents, retail shops, and internet and mobile banking.

Its platform spans three core lines: Payment Aggregator solutions supporting 170+ payment methods for website and app transactions, Biller Network solutions for online bill payments such as utility and credit card bills, and Recurring Payment solutions covering Standing Instructions, eNACH, and UPI AutoPay across Indian bank accounts.

Running that many services meant the customer needed to work with an experienced open source Elastic Stack partner who could design, build, and support a log monitoring environment able to keep pace with the transaction volume.

The Challenge

25+ microservices to troubleshoot

Seamless troubleshooting was needed across more than 25 microservice logs.

No trace ID search

No way to search log messages for a specific trace ID, the unique identifier that follows an API call across microservices.

No per-service latency visibility

Latency monitoring of each individual microservice wasn’t in place.

No HTTP status monitoring

HTTP status codes across services weren’t being tracked centrally.

No restart alerting

No critical alerts were triggered when a microservice restarted.

Needed an experienced Elastic partner

The customer wanted an open source partner with the Elasticsearch expertise to design, build, and support the platform end to end.

The Approach

01
Log ingestion with Logstash
3 Logstash nodes ingest logs from the Payment Aggregator, Biller Network, and Recurring Payment services.
02
Elasticsearch for indexing and search
3 Master nodes and 9 Data nodes handle indexing and search across the full log volume.
03
Kibana for analysis
2 Kibana nodes provide the dashboards used to trace transactions and monitor service health.
04
NGINX+ load balancing
A 2-node NGINX+ layer in the VM environment fronts the stack for the end users accessing it.
05
UAT to production
Ashnik’s Elastic experts built the environment in UAT, tested it, and took it to production with 24×7 post-implementation technical support.
06
Design, deploy, manage, handover
Ashnik’s role covered the full lifecycle: designing the platform, deploying it, managing it in production, and handover.

Architecture

log case study scaled

Scale

METRIC VALUE
Log data ingested ~500 GB/day
Indexing rate ~16,000 events/sec
Events processed ~1.4 billion/day
Log retention 10 days
Logstash nodes 3
Elasticsearch nodes 3 Master + 9 Data
Kibana nodes 2
NGINX+ nodes 2
Microservices monitored 25+

Outcome

Conclusion

Ashnik designed and delivered a log monitoring platform built for the customer’s actual scale, tens of microservices, hundreds of gigabytes of logs a day, and over a billion events processed daily, not a generic ELK deployment.

From UAT through production and into ongoing managed support, the result is a payment gateway platform that can trace any transaction, monitor every microservice, and predict future errors well in advance.