SVP Infrastructure Architect – High-Performance Compute (GPU/CPU) - top hedge fund
Job details
- Employment
- Full-time
- Level
- Executive
- Posted
- Oct 9, 2026
- Last confirmed open
- Oct 9, 2026
About this role
Location
New York City, NY
Industry: Hedge Fund / Investment Management
Employment Type: Permanent, Full-Time
Seniority: Senior / Principal Architect / SVP
A leading global hedge fund is seeking an experienced Infrastructure Architect at SVP level to design, build and scale its next-generation high-performance computing infrastructure.
This is a highly technical and strategically important position, supporting a sophisticated investment environment where large-scale compute capabilities are critical to quantitative research, financial modelling, machine learning, AI and investment strategies.
The successful candidate will take ownership of infrastructure architecture across large-scale GPU and CPU environments, ensuring the organisation can efficiently scale its compute capacity while maintaining exceptional performance, reliability and operational efficiency.
This opportunity will suit an architect with deep expertise in High-Performance Computing (HPC), distributed infrastructure, GPU clusters, large-scale CPU environments and infrastructure automation .
Key Responsibilities
Infrastructure Architecture & Strategy
- Lead the architecture and design of large-scale compute infrastructure supporting quantitative research, AI/ML workloads and investment analytics.
- Develop scalable infrastructure solutions capable of supporting thousands of CPU cores and large GPU clusters.
- Define the technical roadmap for compute infrastructure, balancing performance, scalability, reliability and cost.
- Design hybrid infrastructure architectures spanning on-premises data centres, colocation facilities and public cloud environments.
- Evaluate emerging technologies and recommend infrastructure improvements to support future computational requirements.
- Establish architectural standards and best practices for compute infrastructure.
GPU & CPU Infrastructure
- Architect and scale high-performance GPU environments supporting machine learning, deep learning and complex computational workloads.
- Design and optimise CPU compute clusters supporting large-scale parallel processing and quantitative research.
- Evaluate and integrate NVIDIA GPU technologies, including advanced GPU architectures and high-speed interconnects.
- Develop strategies for efficient GPU allocation, resource scheduling and workload optimisation.
- Ensure compute infrastructure supports demanding performance, latency and throughput requirements.
- Identify and eliminate infrastructure bottlenecks affecting compute-intensive applications.
High-Performance Computing & Distributed Systems
- Design and maintain highly scalable HPC and distributed computing environments.
- Architect workload orchestration and scheduling solutions using technologies such as Slurm, Kubernetes, Ray or equivalent platforms.
- Develop resilient compute architectures capable of handling substantial increases in computational demand.
- Optimise infrastructure performance across networking, storage, memory and compute resources.
- Design solutions for parallel processing, distributed workloads and high-throughput data processing.
Networking, Storage & Infrastructure Performance
- Architect high-bandwidth, low-latency networking solutions supporting GPU and CPU clusters.
- Work with technologies such as InfiniBand, RDMA and high-performance Ethernet.
- Design high-throughput storage solutions capable of supporting large datasets and compute-intensive workloads.
- Evaluate distributed storage platforms and parallel file systems such as Lustre, IBM Spectrum Scale / GPFS or equivalent technologies.
- Ensure infrastructure delivers optimal performance for data-intensive quantitative research.
Automation & Infrastructure Engineering
- Drive infrastructure automation through Infrastructure as Code and modern DevOps practices.
- Develop automated provisioning, deployment and configuration processes for large-scale compute environments.
- Leverage technologies such as Terraform, Ansible, Python and Kubernetes.
- Establish infrastructure monitoring, observability and capacity-planning frameworks.
- Improve infrastructure efficiency through automation, performance tuning and intelligent resource allocation.
Stakeholder Engagement & Technical Leadership
- Partner closely with quantitative researchers, data scientists, software engineers and infrastructure teams.
- Translate complex computational requirements into scalable infrastructure solutions.
- Act as a senior technical authority on GPU, CPU and HPC infrastructure.
- Collaborate with technology leadership on infrastructure investment, capacity planning and long-term architecture.
- Lead technical design reviews and influence infrastructure engineering standards.
Required Experience
- Extensive experience in infrastructure architecture, systems engineering or high-performance computing.
- Proven track record designing and scaling
large enterprise compute environments involving GPU and CPU clusters
.
- Strong understanding of HPC architecture, distributed systems and parallel computing.
- Hands-on expertise with Linux-based infrastructure and large-scale server environments.
- Experience architecting GPU infrastructure, ideally using NVIDIA technologies.
- Strong knowledge of compute orchestration, workload scheduling and cluster management.
- Experience with Kubernetes, Slurm or comparable compute scheduling platforms.
- Deep understanding of high-performance networking, including InfiniBand, RDMA or low-latency Ethernet.
- Experience designing scalable storage architectures for compute-intensive workloads.
- Strong infrastructure automation skills using Python, Terraform, Ansible or similar technologies.
- Experience with infrastructure capacity planning, performance optimisation and cost management.
- Ability to communicate complex technical architecture to senior engineering and business stakeholders.
Highly Desirable Experience
- Previous experience within a
hedge fund, quantitative trading firm, proprietary trading organisation or investment bank
.
- Experience supporting quantitative research, algorithmic trading or financial modelling environments.
- Experience designing large-scale AI/ML infrastructure and GPU compute platforms.
- Familiarity with NVIDIA CUDA, NVIDIA AI Enterprise or NVIDIA networking technologies.
- Experience with distributed computing frameworks such as Ray, Dask or Apache Spark.
- Knowledge of cloud-based HPC solutions across AWS, Azure or Google Cloud.
- Experience scaling infrastructure across multiple data centres or geographic regions.
- Familiarity with containerised HPC workloads and GPU-aware Kubernetes scheduling.
- Experience evaluating infrastructure performance, power consumption and total cost of ownership.
For more info you can contact [email protected]