AI Solution Architect

Posted Sep 18

This is a fully remote position, open to applicants in India.

📋 Description

• Take ownership of the complete architecture for AI Factory and enterprise AI solutions, from initial requirements to production readiness.

• Evaluate AI/ML workload needs for training, fine-tuning, inference, batch processing, and high-performance computing.

• Architect GPU compute systems, including NVIDIA HGX/DGX/OEM platforms, multi-GPU configurations, NVLink/NVSwitch, and GPU resource management.

• Create high-performance AI networking solutions utilizing 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.

• Develop AI storage and data architectures leveraging object storage and parallel file systems such as Ceph and WEKA.

• Specify the architecture for AI platforms across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.

• Set architecture standards encompassing security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.

• Generate reference architectures, high-level/low-level designs, capacity models, bills of materials, technology assessments, and implementation roadmaps.

• Lead technical evaluations, proof-of-concept initiatives, vendor assessments, and architecture review boards.

• Collaborate with teams across infrastructure, networking, security, storage, cloud, data, application, and operations.

• Establish performance, availability, scalability, security, and cost goals, validating architecture against measurable acceptance criteria.

• Provide technical guidance throughout deployment, migration, integration, troubleshooting, and the transition to production.

• Create AI Factory reference architectures, solution blueprints, high-level designs (HLDs), low-level designs (LLDs), architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, bills of materials, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks, and acceptance criteria.


⛳️ Requirements

• Over 10 years of experience in infrastructure, cloud, enterprise architecture, or solution architecture, with a strong background in AI/GPU infrastructure.

• Demonstrated experience in designing large-scale enterprise platforms and converting business requirements into technical architectures.

• Practical knowledge of physical infrastructure, GPU systems, networking, storage, and Linux platforms.

• A Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field is preferred.

• Preferred qualifications include NVIDIA certifications or equivalent credentials in GPU/AI infrastructure.

• AWS Solutions Architect or Azure Solutions Architect certification is preferred.

• TOGAF or an equivalent enterprise architecture certification is preferred.

• CCNP/CCIE or equivalent networking certification is preferred.

• CISSP or an equivalent security certification is preferred.

• Kubernetes certifications such as CKA/CKAD are preferred.

• Red Hat/Linux certifications are preferred.

• Expertise in AI/ML architecture, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator, and AI platform ecosystems.

• Familiarity with PyTorch, TensorFlow, and JAX, including operational knowledge of training and inference workloads.

• Understanding of GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization, and workload placement.

• Knowledge of LLM, generative AI, RAG, fine-tuning, model serving, and inference architecture.

• Experience with NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology, and multi-GPU performance architecture.

• Knowledge of DGX/HGX/OEM GPU server architecture and lifecycle management.

• Understanding of AI Factory capacity planning, rack density, power, cooling, commissioning, and lifecycle strategies.

• Proficiency in 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, and QoS.

• Familiarity with NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum, and BlueField DPU technologies.

• Experience with GPU east-west traffic, GPUDirect RDMA, and network performance troubleshooting.

• Knowledge of parallel file systems, object storage, NFS, NVMe/NVMe-oF, and high-throughput data pipelines.

• Familiarity with technologies such as Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp, or equivalent.

• Understanding of data lake/lakehouse concepts, metadata, lineage, data movement, and data lifecycle.

• Skills in GPUDirect Storage and storage/network performance optimization.

• Proficiency in Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling, and HPC or equivalent workload schedulers.

• Experience with model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management, and platform integration.

• Familiarity with AWS and/or Azure AI infrastructure and security services.

• Knowledge of hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning, and FinOps.

• Understanding of Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, and security considerations for GPU/DPU/container/Kubernetes/firmware/supply-chain.

• Skills in encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation, and secure model access.

• Familiarity with Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM, and infrastructure telemetry.

• Experience in monitoring across GPU, CPU, memory, network, storage, power, and thermal domains.

• Knowledge of high availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design, and capacity forecasting.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Generous paid time off and flexible working arrangements.

• Opportunities for professional development and continuous learning.

• Collaborative and innovative work environment.

People also viewed

PortX1 day ago

Solutions Architect – Enterprise Integration

US flagTexas OnlyFull-timeSolutions Engineer$180k – $200k/year
ApplyView job
PortX1 day ago

Solutions Architect – Enterprise Integration

US flagUnited States OnlyFull-timeSolutions Engineer$170k – $200k/year
ApplyView job
PortX1 day ago

Solutions Architect – Enterprise Integration

US flagTexas OnlyFull-timeSolutions Engineer$180k – $200k/year
ApplyView job
SailPoint1 day ago

Solution Engineer

US flagIllinois OnlyFull-timeSolutions Engineer$73.1k – $123.2k/year
ApplyView job
SailPoint1 day ago

Advisory Solutions Consultant

US flagNorth Carolina OnlyFull-timeSolutions Engineer$116.9k – $197.1k/year
ApplyView job
Squirro1 day ago

Solutions Architect

EuropeFull-timeSolutions Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers