Powered by pgvector · cosine kNN
No similar jobs yet — embeddings have not been generated for this listing. Run npm run seed:embeddings with AI Gateway access (AI_GATEWAY_API_KEY) configured.
Synechron
We are seeking an experienced Platform Engineer to build, operate, and continuously improve cloud-native platforms hosted on AWS. The ideal candidate will have strong expertise in AWS infrastructure, containerized wor…
Your match
See how you fit
Scored against this job in seconds
Your account
Sign in to apply
Your profile and your match for this job appear right here.
sign in above to apply · via LinkedIn
About the role
We are seeking an experienced Platform Engineer to build, operate, and continuously improve cloud-native platforms hosted on AWS. The ideal candidate will have strong expertise in AWS infrastructure, containerized workloads, CI/CD automation, observability, and production support.This role requires hands-on experience managing Amazon ECS environments, troubleshooting complex production incidents, implementing Infrastructure as Code (IaC), and driving platform reliability, scalability, and automation initiatives. The successful candidate will work closely with engineering teams to ensure high availability, performance, security, and operational excellence across enterprise platforms.Key ResponsibilitiesDesign, implement, and maintain highly available, scalable, and secure AWS-based platform infrastructure.Manage and support containerized applications running on Amazon ECS (EC2 and/or Fargate).Monitor platform health and respond to production incidents, ensuring minimal downtime and rapid resolution.Troubleshoot application performance issues, latency bottlenecks, infrastructure failures, and deployment-related incidents.Optimize ECS clusters, Auto Scaling Groups, load balancers, and cloud resources for performance and cost efficiency.Develop and maintain Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar technologies.Build and enhance CI/CD pipelines to support automated deployment, testing, and release processes.Configure and maintain monitoring, logging, alerting, and observability platforms to improve operational visibility.Perform Root Cause Analysis (RCA) and implement preventative measures to improve platform stability.Collaborate with software engineering teams to enhance application reliability, deployment practices, and operational readiness.Implement cloud security, governance, compliance, and operational best practices.Drive platform automation initiatives to reduce manual intervention and improve operational efficiency.Required Skills & ExperienceCloud & Platform Engineering5+ years of experience as a Platform Engineer, DevOps Engineer, Cloud Engineer, or Site Reliability Engineer (SRE).Strong hands-on experience with AWS services including:Amazon ECSEC2Application Load Balancer (ALB)Network Load Balancer (NLB)CloudWatchAuto ScalingIAMVPC and NetworkingContainers & OrchestrationStrong experience managing containerized applications using:DockerAmazon ECSKubernetes (desirable)Production Support & TroubleshootingProven experience investigating and resolving production incidents related to:High latency5xx application errorsApplication crashesMemory leaksResource bottlenecksCapacity and scaling issuesStrong understanding of:Application performance tuningResource optimizationContainer performance managementInfrastructure troubleshootingInfrastructure & AutomationExpertise with Infrastructure as Code (IaC):TerraformCloudFormationStrong scripting and automation skills using:PythonBashPowerShellDevOps & CI/CDExperience implementing and supporting CI/CD pipelines using:GitHub ActionsGitLab CI/CDJenkinsAzure DevOpsMonitoring & ObservabilityHands-on experience with monitoring and observability platforms including:CloudWatchDatadogGrafanaSplunkPrometheusELK StackDesired ExperienceStrong experience supporting production ECS workloads and container lifecycle management.Knowledge of Out-of-Memory (OOM) analysis, JVM tuning, and container resource optimization.Experience implementing Blue-Green and Canary deployment strategies.Exposure to Site Reliability Engineering (SRE) principles and operational best practices.Experience implementing autoscaling strategies using CPU, memory, throughput, and application-level metrics.Understanding of FinOps practices and cloud cost optimization.Experience with incident management, post-incident reviews, and operational excellence frameworks.
sign in above to apply · via LinkedIn
Your job hunt, handled
Ask about any role and get a straight answer on your fit. Then stop searching: new matches land in your WhatsApp the moment they’re listed.
Free for jobseekers