WhatsApp Telegram LinkedIn

Site Reliability Engineer (Observability) :: New York, NY

Contract

C2C contract jobs

Job Description — Site Reliability Engineer (Observability)

Job Location: New York City, NY (Hybrid — Onsite Required)

Job Type: Long-Term Contract

Client – Gspann / End client not disclosed at this moment. 

 

Key Skills: Retail domain with skills like Splunk, LogicMonitor/Dynatrace/Datadog tool capabilities.

 

 

Role Summary

We are seeking an experienced Site Reliability Engineer (SRE) with strong Observability expertise to support large-scale, customer-facing retail platforms. The ideal candidate has hands-on experience with enterprise monitoring toolchains and a track record of ensuring system reliability and uptime in high-traffic retail/eCommerce environments. This is a hybrid, long-term contract role based in NYC with required onsite attendance.

 

Key Responsibilities

Design and maintain observability pipelines covering logs, metrics, traces, and synthetic monitoring across production and non-production environments.

Build dashboards, alerts, and SLO/SLI frameworks using Splunk and one or more of LogicMonitor, Dynatrace, or Datadog.

Partner with application, infrastructure, and platform teams to define monitoring coverage for retail-critical systems (order management, checkout, inventory, POS, fulfillment).

Lead incident response and root cause analysis, using observability data to reduce MTTD and MTTR.

Establish proactive alerting strategies that reduce noise while maintaining signal fidelity for critical services.

Support high-traffic seasonal events (holiday peak, promotions, flash sales) with readiness reviews and real-time monitoring support.

Automate observability configuration (monitoring-as-code) and integrate observability into CI/CD pipelines.

Document runbooks, escalation paths, and post-incident reviews.

 

Required Skills & Experience

7+ years in SRE, DevOps, or Observability/Monitoring engineering roles.

Hands-on expertise with Splunk (search, dashboards, alerting, log pipeline management).

Practical experience with at least one of: LogicMonitor, Dynatrace, or Datadog.

Prior experience supporting retail or eCommerce platforms, ideally including peak-load event support.

Strong understanding of distributed systems, microservices, and cloud infrastructure (AWS/Azure/GCP).

Experience with incident management, on-call rotations, and postmortem/RCA processes.

Scripting proficiency (Python, Shell, or similar) for monitoring/alerting automation.

Familiarity with containerized environments (Kubernetes, Docker) and CI/CD tooling.

Solid understanding of SLIs, SLOs, error budgets, and reliability engineering principles.

 

Preferred

Additional observability tools (New Relic, Grafana, Prometheus, ELK).

APM instrumentation experience (OpenTelemetry, tracing frameworks).

Infrastructure-as-code experience (Terraform, Ansible).

Familiarity with retail systems: POS, OMS, WMS, inventory/fulfillment platforms.

Relevant cloud or observability tool certifications.

To apply for this job email your details to pratiksha.hatkar@nytpcorp.com

×

Post your C2C job instantly

Quick & easy posting in 10 seconds

Keep it concise - you can add details later
Please use your company/professional email address
Simple math question to prevent spam