HomeJob Openings

Job Openings

About BerryBytes

BerryBytes is a DevOps and cloud services partner helping clients modernize their infrastructure, adopt AI-driven workflows, and unlock data-driven decision-making. We work across Azure, AWS, and modern data platforms to deliver scalable, secure solutions for enterprise clients. This engagement is part of our growing data analytics practice.

Role Overview

We are looking for an experienced Microsoft Fabric & Data Platform Consultant to join a client-facing project. You will design, build, and optimize end-to-end data solutions using Microsoft Fabric, Azure data services, and Tableau. The ideal candidate brings hands-on production experience — not just platform familiarity — and can work independently with minimal oversight.

Key Responsibilities

  • Design and implement data pipelines using Microsoft Fabric (Data Factory, Lakehouse, Warehouse, OneLake)
  • Build and optimize data models, semantic layers, and Direct Lake Power BI integrations
  • Architect and manage Azure data services including Azure Data Lake Storage, Azure Synapse, and Azure Data Factory
  • Develop interactive Tableau dashboards and reports connected to Fabric and Azure data sources
  • Define and enforce data governance, workspace management, and security policies within Fabric
  • Collaborate with client stakeholders to gather requirements and translate them into scalable data architecture
  • Conduct performance tuning and troubleshooting across pipelines, queries, and reporting layers
  • Produce technical documentation — architecture diagrams, data flow maps, and runbooks
  • Support data migration from legacy platforms (on-prem SQL, Azure SQL DW) to Fabric Lakehouse Warehouse

Required Skills & Experience

Microsoft Fabric (must-have):

  • Hands-on production experience with Microsoft Fabric — OneLake, Lakehouse, Data Warehouse, Notebooks, and Pipelines
  • Proficiency in Direct Lake mode for Power BI and multi-workspace governance
  • Experience with Fabric security models, capacity management, and workspace admin

Azure Data Platform:

  • Strong working knowledge of Azure Data Lake Storage Gen2, Azure Data Factory, and Azure Synapse Analytics
  • Experience with Azure DevOps pipelines for CI/CD of data assets
  • Familiarity with Azure Entra ID (AAD) and role-based access control for data resources

Tableau:

  • Proficient in Tableau Desktop and Tableau Server/Cloud for enterprise reporting
  • Ability to connect Tableau to Fabric and Azure SQL sources with performant live or extract connections
  • Dashboard design skills with a focus on business usability

General:

  • Strong SQL skills (T-SQL and/or Spark SQL)
  • Experience with Python or PySpark for data transformation in notebooks
  • Ability to document architectures clearly and communicate with non-technical stakeholders

Nice to Have

  • Microsoft Certified: Fabric Analytics Engineer Associate (DP-600) or equivalent
  • Experience with Tableau Prep for data preparation workflows
  • Familiarity with dbt (data build tool) within Fabric or Azure environments
  • Exposure to LLM/AI integrations on Azure (Azure OpenAI, Copilot in Fabric)
  • Prior experience with DevOps consulting engagements or client-facing delivery

About the Role

We are a fast-growing AI infrastructure company building cutting-edge GPU cloud platforms and high-performance inference solutions that empower AI developers, startups, and enterprises worldwide. As we scale our global operations, we are looking for a skilled and hands-on AI Infra Engineer – SRE (Kubernetes) to join our Global Infrastructure team.

Role Overview

This is a critical hands-on position focused on the reliability, performance, and operational excellence of large-scale, high-performance AI/ML GPU clusters in our data centers. As an AI Infra Engineer – SRE (Kubernetes), you will design, operate, and optimize Kubernetes-based infrastructure to ensure maximum uptime, efficiency, and scalability for demanding AI workloads. You will bring deep expertise in system-level troubleshooting, GPU cluster management, and automation to keep our platforms running at peak performance.

Key Responsibilities

  • Design, build, and maintain scalable, production-grade AI/ML infrastructure using Kubernetes.
  • Proactively monitor GPU cluster health, performance, and utilization across compute, accelerators, storage, and networking layers, performing root-cause analysis and resolution.
  • Develop and implement automation for infrastructure provisioning, configuration, and ongoing management.
  • Own the complete GPU node lifecycle — including provisioning, dynamic scaling, maintenance, decommissioning, and zero-downtime upgrades of GPU-enabled nodes in Kubernetes environments.
  • Build and improve CI/CD pipelines for reliable infrastructure deployment and orchestration.
  • Enforce security best practices, compliance standards, and operational excellence across the infrastructure stack.
  • Lead incident response and post-incident improvements for issues related to GPUs, CPUs, high-speed storage, and networks.
  • Manage end-to-end customer GPU resource provisioning — from request intake and configuration to onboarding, troubleshooting, and support — ensuring high levels of customer satisfaction.
  • Stay up to date with the latest GPU hardware, software, and orchestration technologies, integrating relevant advancements into our platforms.
  • Be available for occasional regional or international travel to data center locations as required.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of practical experience in data center operations, infrastructure engineering, or site reliability engineering.
  • Strong background in infrastructure automation using tools such as Terraform and Ansible.
  • Deep hands-on experience with Kubernetes in large-scale environments, including:
    • NVIDIA GPU Operator for GPU driver management, device plugins, container toolkit, and monitoring (DCGM).
    • NVIDIA Network Operator for high-performance networking, RDMA, and GPUDirect support.
    • CNI (Container Network Interface) and CSI (Container Storage Interface) plugins tailored for AI/ML workloads.
    • Integration with job schedulers such as Slurm in Kubernetes clusters.
  • Proficiency in Linux system administration and scripting (Python, Bash).
  • Experience with observability stacks including Prometheus, Grafana, and Loki.
  • Solid understanding of GPU architecture, NVIDIA CUDA, NCCL, and AI/ML frameworks is a strong plus.
  • Excellent troubleshooting skills with the ability to analyze complex system logs and performance metrics.
  • Strong communication and collaboration skills to work effectively with engineering and operations teams.

Reports To : Director of Product Development

Why You’ll Love Working With Us

This isn’t just another AI job. You’ll be part of a pioneering team pushing the boundaries of what autonomous AI systems can do for Platform Engineering. You’ll have the freedom to innovate, the resources to build at scale, and the support of a collaborative, forward-thinking environment.

We are seeking a skilled d AI Agentic Engineer to join our innovative team. The ideal candidate will have a strong background in artificial intelligence, natural language processing, and software development. This role involves designing, developing, and implementing advanced RAG systems and Agentic AI based solutions that enhance user interaction and content retrieval.

What We’re Looking For

  • You eat, sleep, and breathe generative AI — prompt and context engineering, retrieval-augmented generation (RAG), and autonomous agents are your playground.
  • You know your way around modern AI agent frameworks like LangChain, LlamaIndex, Semantic Kernel, crewAI, and AutoGen — and you’re excited to push their limits.
  • Vector databases like Pinecone, Weaviate, or Chroma? You’re comfortable querying and managing them to power semantic search.
  • Full-stack skills? Absolutely. React + TypeScript on the frontend, Node.js or Python microservices on the backend, and REST or gRPC APIs.
  • DevOps savvy: Kubernetes, Terraform or AWS CDK, plus monitoring tools like Grafana and Prometheus are in your toolkit.

Key Responsibilities

  • Design and Build AI Agents: Create autonomous agents that can reason, plan, act, and collaborate.
  • Prompt Engineering: Develop advanced prompts and roles to guide agent behavior.
  • Memory & Context Integration: Implement systems for agents to store, recall, and use memory in conversations.
  • Tool & API Integration: Enable agents to use external tools and APIs for real-world automation.
  • Multi-Step Reasoning: Build agents capable of decomposing tasks, planning, and self-correcting.
  • Multi-Agent Systems: Deploy and manage groups of agents that collaborate and delegate.
  • Ecosystem Automation: Launch and maintain agentic systems that automate business and technical workflows.
  • Collaborate with data scientists and engineers to refine algorithms and improve the performance of AI models.
  • Conduct thorough testing and validation of developed systems to ensure accuracy and reliability.
  • Stay updated with industry trends and advancements in AI, machine learning, and natural language processing.

The Tech We Love

  • AI Agent & Orchestration: LangChain, LangGraph, CrewAU, LlamaIndex, Semantic Kernel, AutoGen
  • Protocol: MCP, Agent2Agent
  • Vector DBs: Pinecone, Weaviate, Chroma
  • Observability & Evaluation: LangSmith, Helicone, PromptLayer, RAGAS
  • CI/CD for LLMs: PromptOps, LlamaTest, GitHub Actions with AI evaluation workflows
  • Telephony: Twilio Programmable Voice, SIP, VAPI

If you are passionate about advancing AI technologies and have the skills to build innovative RAG and agentic systems, we encourage you to apply. Join us in shaping the future of intelligent applications!

Reports To: Director of Cloud Infrastructure

Role Overview:
Join our Core Kubernetes Operator Development team, where we’re pushing the boundaries of Kubernetes innovation. As a Kubernetes Controller Developer (Golang), you will play a crucial role in building “01”, our cloud-agnostic Platform as a Service (PaaS), driven by full-fledged Kubernetes operators and agents.

This position requires a strong background in Kubernetes internals and Golang programming, particularly in developing and managing Kubernetes controllers. If you’re a proactive problem solver with experience in building cloud-native infrastructure, this is your opportunity to contribute to a transformative platform.

We highly encourage candidates with a solid programming foundation and a hunger to explore the cloud-native world to apply. Comprehensive onboarding and professional development support will be provided.

Key Responsibilities (Not limited to):

  • Collaborate in Agile teams, taking ownership of development stories with minimal supervision.
  • Partner with internal teams and clients to accurately capture technical requirements.
  • Design, build, deploy, and maintain Kubernetes controllers and operators using Golang.
  • Identify gaps in current systems and propose or implement technical improvements.
  • Apply best practices across the full software development lifecycle.
  • Create and execute unit, regression, and E2E tests for operator reliability.
  • Work in Linux environments and troubleshoot issues in containerized applications.
  • Contribute to CI/CD workflows for seamless testing and deployment.

Essential Skillset:

  • Kubernetes Controller Development: Proven expertise in building and maintaining controllers and operators.
  • Proficiency in Golang: 2+ years writing idiomatic, well-tested Go code for Kubernetes projects.
  • Deep understanding of Kubernetes APIs and libraries including client-go, CRDs, and API extensions.
  • Hands-on experience with:
    • Kubebuilder – For scaffolding controllers and CRDs
    • Operator SDK – For building Operators with OLM support
    • controller-runtime – For abstracting Kubernetes client logic
  • Strong testing skills, including unit, load, and E2E tests for operators.
  • Familiarity with containerization (Docker) and orchestration (Kubernetes).
  • Comfortable working in Linux with debugging tools and CLI.
  • 2+ years experience working with CI/CD tools like Jenkins, GitHub Actions, Tekton, or similar.

Preferred Skills (Nice to Have):

  • CKA or CKAD certifications.
  • Hands-on experience managing production-grade Kubernetes clusters.
  • Knowledge of Infrastructure as Code tools (e.g., Terraform).
  • Exposure to major cloud providers: AWS, GCP, or Azure.
  • Scripting experience in Shell or Python.

What We Offer:

  • A chance to build infrastructure automation tools that power real-world workloads.
  • Opportunity to work on bleeding-edge cloud-native technologies with a global impact.
  • Collaborative and innovation-driven culture, with strong engineering mentorship.
  • Remote-friendly setup and flexible work culture.
  • Career development in one of the most in-demand areas of DevOps.

Get the latest BerryBytes updates by subscribing to our Newsletter!

Enterprise AI Acceleration Unleashed

Copyright © 2026 BerryBytes. All Rights Reserved.