Start Your Search Here

push notification bell

Would you like to receive notifications about jobs in Santa Clara?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

FlexAI

Santa Clara / Global

Senior Backend Engineer

Job Description

Senior Backend Engineer (Infrastructure & AI Platform)FlexAI is looking for a Senior Backend Engineer (Infrastructure & AI Platform) with deep Golang expertise to architect and build the core backend systems powering our next-generation AI compute and PaaS platform. This role sits at the intersection of distributed systems, cloud infrastructure, and AI platform engineering — enabling large-scale model training, inference, and orchestration across heterogeneous compute. This is not a traditional backend role; you will be building platform-grade systems that support AI runtimes, scheduling, resource orchestration, and multi-tenant cloud infrastructure.As a Senior Backend Engineer, you'll drive backend architecture, scale platform services, and build high-performance infrastructure components that power AI workloads in production environments — influencing how the platform evolves from Beta to enterprise-grade deployment. Expect high ownership and technical autonomy in a research-driven, deep-tech environment — not SaaS CRUD apps.This position is In-Person and located at our San Jose, CA Office.Core Platform & Infrastructure Backend:Architect and develop high-performance Golang services for FlexAI's AI PaaS and infrastructure platformBuild internal APIs powering model deployment, job scheduling, and compute lifecycle managementDevelop components interfacing with GPU/compute infrastructure and AI runtimesDistributed Systems & Scalability:Design and scale microservices and event-driven systems for high-throughput AI workloadsOptimize for low latency, high concurrency, and fault toleranceImplement service-to-service communication (gRPC/REST, message queues, async pipelines)Drive reliability, observability, and resilience across servicesAI Platform Integration:Collaborate with AI/ML and Runtime teams to integrate systems with training pipelines, inference infrastructure, experimentation workflows, and dataset/artifact managementEnable orchestration across cloud and on-prem environmentsBuild abstractions that simplify AI infrastructure consumptionCloud-Native & Platform Engineering:Design cloud-native, Kubernetes-native servicesWork with DevOps/SRE on CI/CD, deployment automation, and scalabilityContribute to architecture decisions for multi-region, multi-cloud infrastructureImprove monitoring, logging, and diagnosticsTechnical Leadership:Lead architecture reviews and set engineering standardsMentor engineers and guide complex problem-solvingDrive long-term roadmap for backend infrastructure and AI platform capabilitiesPartner with Product, Runtime, and Infra leadership to translate requirements into scalable systemsTech Stack (Indicative):Languages: Golang (Primary), Python (Secondary)Infrastructure: Kubernetes, Docker, Cloud (AWS/GCP/Azure)Architecture: Microservices, gRPC, Event-driven systemsData: SQL + NoSQL databases, caching, streaming systemsObservability: Prometheus, Grafana, OpenTelemetry (or similar)What You'll Need to Be SuccessfulCore Engineering:5+ years of Backend or Infrastructure Engineering experienceExpert-level proficiency in Golang (must-have, heavy hands-on)Strong experience building production-grade distributed systemsProven track record on infrastructure platforms, PaaS, or deep-tech systemsInfrastructure & Systems:Deep understanding of cloud-native architectures and containerized environmentsStrong experience with Kubernetes, Docker, and cluster orchestrationFamiliarity with compute scheduling, resource management, or platform runtimes is a strong plusDatabases & Data Systems:Experience with distributed databases (PostgreSQL, Cassandra, DynamoDB, etc.)Strong understanding of caching, queues, and streaming systems (Redis, Kafka, etc.)AI / Platform Exposure (Highly Preferred):Experience on AI/ML platforms, model infrastructure, or data platformsFamiliarity with ML pipelines, inference systems, or GPU-backed workloadsExposure to PyTorch, TensorFlow infrastructure, or model serving systems is a plusIdeal Candidate Profile (Who Will Thrive Here)Infra-first backend engineers (not just API developers)Background in AI infra, cloud platforms, developer platforms, or deep-tech systemsStrong systems thinkers who enjoy low-level performance, scalability, and architecture challengesStartup-minded builders comfortable in ambiguous, high-ownership environmentsWhat We OfferCompetitive salary and benefits packageWork on cutting-edge AI infrastructureBuild products used by developers and enterprisesHigh ownership, fast execution, real impactCollaborative, high-caliber team
Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location