Founding Engineer, AI Infra

Goaly AI

📍 San Francisco, California, United States, UN0💼 Tempo pieno🕐 5 giorni fa

Crea un account gratis in 30 secondi: ottieni anche il match score AI con il tuo CV.

Descrizione

About Goaly At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog. Role Overview You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale. Key Responsibilities Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton). Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability. Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding. Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration. Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems. Requirements 5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models. Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies. Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling. Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production. Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems. Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry). Bonus Points Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models. Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects. Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems. Why join us? Solve Uncharted Scaling Problems: Move beyond incremental improvements. You'll architect systems that enable 90% of frontier model performance at 10% of the cost, tackling the most challenging problems in distributed compute for LLMs and SLMs—a problem space few have genuinely cracked. Redefine AI Infrastructure: Leverage your deep expertise to fundamentally reshape how AI models are trained and deployed globally. This isn't just optimization; it's about building an entirely new paradigm that democratizes access to SOTA AI. Direct Impact on Business & Science: Your work directly underpins both our enterprise solutions and our open-source contributions. See your infrastructure choices enable breakthroughs for diverse industries and advance the entire field. Ownership & Autonomy in a Flat Structure: Work in a small, elite team where your voice is paramount. You'll have complete ownership over critical infrastructure components, setting technical strategy with direct access to founding FAANG-level expertise. Proven, Cutting-Edge Edge: We've already demonstrated 4-10x speedups for a few frontier models and are pushing limits on next-gen models. You'll be working with a team that has a track record of delivering breakthrough performance, not just promises. Meaningful Equity: Receive an early-stage package with massive upside—you directly capture the value of the infrastructure breakthroughs you create.

Candidati ora →

TalentyGo è un aggregatore di offerte da fonti pubbliche. Verifica sempre le informazioni direttamente con l'azienda. La candidatura avviene tramite il sito originale dell'azienda; TalentyGo non gestisce processi di selezione.