
C2C contract jobs
Job Description — Site Reliability Engineer (Observability)
Job Location: New York City, NY (Hybrid — Onsite Required)
Job Type: Long-Term Contract
Client – Gspann / End client not disclosed at this moment.
Key Skills: Retail domain with skills like Splunk, LogicMonitor/Dynatrace/Datadog tool capabilities.
Role Summary
We are seeking an experienced Site Reliability Engineer (SRE) with strong Observability expertise to support large-scale, customer-facing retail platforms. The ideal candidate has hands-on experience with enterprise monitoring toolchains and a track record of ensuring system reliability and uptime in high-traffic retail/eCommerce environments. This is a hybrid, long-term contract role based in NYC with required onsite attendance.
Key Responsibilities
Design and maintain observability pipelines covering logs, metrics, traces, and synthetic monitoring across production and non-production environments.
Build dashboards, alerts, and SLO/SLI frameworks using Splunk and one or more of LogicMonitor, Dynatrace, or Datadog.
Partner with application, infrastructure, and platform teams to define monitoring coverage for retail-critical systems (order management, checkout, inventory, POS, fulfillment).
Lead incident response and root cause analysis, using observability data to reduce MTTD and MTTR.
Establish proactive alerting strategies that reduce noise while maintaining signal fidelity for critical services.
Support high-traffic seasonal events (holiday peak, promotions, flash sales) with readiness reviews and real-time monitoring support.
Automate observability configuration (monitoring-as-code) and integrate observability into CI/CD pipelines.
Document runbooks, escalation paths, and post-incident reviews.
Required Skills & Experience
7+ years in SRE, DevOps, or Observability/Monitoring engineering roles.
Hands-on expertise with Splunk (search, dashboards, alerting, log pipeline management).
Practical experience with at least one of: LogicMonitor, Dynatrace, or Datadog.
Prior experience supporting retail or eCommerce platforms, ideally including peak-load event support.
Strong understanding of distributed systems, microservices, and cloud infrastructure (AWS/Azure/GCP).
Experience with incident management, on-call rotations, and postmortem/RCA processes.
Scripting proficiency (Python, Shell, or similar) for monitoring/alerting automation.
Familiarity with containerized environments (Kubernetes, Docker) and CI/CD tooling.
Solid understanding of SLIs, SLOs, error budgets, and reliability engineering principles.
Preferred
Additional observability tools (New Relic, Grafana, Prometheus, ELK).
APM instrumentation experience (OpenTelemetry, tracing frameworks).
Infrastructure-as-code experience (Terraform, Ansible).
Familiarity with retail systems: POS, OMS, WMS, inventory/fulfillment platforms.
Relevant cloud or observability tool certifications.
To apply for this job email your details to pratiksha.hatkar@nytpcorp.com