
AI Solution Architect
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in India.
• Take ownership of the complete architecture for AI Factory and enterprise AI solutions, from initial requirements to production readiness.
• Evaluate AI/ML workload needs for training, fine-tuning, inference, batch processing, and high-performance computing.
• Architect GPU compute systems, including NVIDIA HGX/DGX/OEM platforms, multi-GPU configurations, NVLink/NVSwitch, and GPU resource management.
• Create high-performance AI networking solutions utilizing 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
• Develop AI storage and data architectures leveraging object storage and parallel file systems such as Ceph and WEKA.
• Specify the architecture for AI platforms across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
• Set architecture standards encompassing security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
• Generate reference architectures, high-level/low-level designs, capacity models, bills of materials, technology assessments, and implementation roadmaps.
• Lead technical evaluations, proof-of-concept initiatives, vendor assessments, and architecture review boards.
• Collaborate with teams across infrastructure, networking, security, storage, cloud, data, application, and operations.
• Establish performance, availability, scalability, security, and cost goals, validating architecture against measurable acceptance criteria.
• Provide technical guidance throughout deployment, migration, integration, troubleshooting, and the transition to production.
• Create AI Factory reference architectures, solution blueprints, high-level designs (HLDs), low-level designs (LLDs), architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, bills of materials, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks, and acceptance criteria.
• Over 10 years of experience in infrastructure, cloud, enterprise architecture, or solution architecture, with a strong background in AI/GPU infrastructure.
• Demonstrated experience in designing large-scale enterprise platforms and converting business requirements into technical architectures.
• Practical knowledge of physical infrastructure, GPU systems, networking, storage, and Linux platforms.
• A Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field is preferred.
• Preferred qualifications include NVIDIA certifications or equivalent credentials in GPU/AI infrastructure.
• AWS Solutions Architect or Azure Solutions Architect certification is preferred.
• TOGAF or an equivalent enterprise architecture certification is preferred.
• CCNP/CCIE or equivalent networking certification is preferred.
• CISSP or an equivalent security certification is preferred.
• Kubernetes certifications such as CKA/CKAD are preferred.
• Red Hat/Linux certifications are preferred.
• Expertise in AI/ML architecture, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator, and AI platform ecosystems.
• Familiarity with PyTorch, TensorFlow, and JAX, including operational knowledge of training and inference workloads.
• Understanding of GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization, and workload placement.
• Knowledge of LLM, generative AI, RAG, fine-tuning, model serving, and inference architecture.
• Experience with NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology, and multi-GPU performance architecture.
• Knowledge of DGX/HGX/OEM GPU server architecture and lifecycle management.
• Understanding of AI Factory capacity planning, rack density, power, cooling, commissioning, and lifecycle strategies.
• Proficiency in 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, and QoS.
• Familiarity with NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum, and BlueField DPU technologies.
• Experience with GPU east-west traffic, GPUDirect RDMA, and network performance troubleshooting.
• Knowledge of parallel file systems, object storage, NFS, NVMe/NVMe-oF, and high-throughput data pipelines.
• Familiarity with technologies such as Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp, or equivalent.
• Understanding of data lake/lakehouse concepts, metadata, lineage, data movement, and data lifecycle.
• Skills in GPUDirect Storage and storage/network performance optimization.
• Proficiency in Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling, and HPC or equivalent workload schedulers.
• Experience with model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management, and platform integration.
• Familiarity with AWS and/or Azure AI infrastructure and security services.
• Knowledge of hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning, and FinOps.
• Understanding of Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, and security considerations for GPU/DPU/container/Kubernetes/firmware/supply-chain.
• Skills in encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation, and secure model access.
• Familiarity with Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM, and infrastructure telemetry.
• Experience in monitoring across GPU, CPU, memory, network, storage, power, and thermal domains.
• Knowledge of high availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design, and capacity forecasting.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and flexible working arrangements.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
PortX
PortX
PortX
SailPoint
Get handpicked remote jobs straight to your inbox weekly.