# Omnistrate Documentation > Build your software distribution channels in days, not years. Comprehensive documentation for Omnistrate's agentic control plane generator. # Introduction # Welcome to Omnistrate ## Introduction to Omnistrate **Build your software distribution channels—in days, not years!** Welcome to the comprehensive documentation for Omnistrate, the world's first agentic control plane generator. We help ISVs transform their software into multi-cloud SaaS products with a private control plane generated for their product, domain, and customer experience. Whether you're a developer, DevOps engineer, platform team member, security professional, or FinOps specialist, this documentation provides the guidance you need to leverage Omnistrate's capabilities for your specific role and responsibilities. ## How Omnistrate Transforms Your Product Distribution Omnistrate's primary strength lies in helping you distribute your software across multiple channels with integrated billing and complete automation. No matter where your customers want to deploy—whether in their own cloud accounts, on-premises, or as a managed service—we make it possible in days instead of years. Your generated control plane automates all distribution channels: - **On-Premises & Air-Gapped** - Deploy in isolated environments with complete control - **BYOC (Bring Your Own Cloud)** - Run in customer cloud accounts with full isolation - **PaaS (Platform as a Service)** - Offer managed platform experiences - **SaaS (Software as a Service)** - Provide fully managed cloud services - **AaaS (Agent as a Service)** - Distribute and deploy AI agents in any environment with automated scaling, intelligent routing, and seamless integration capabilities The result? You can focus on your core product while Omnistrate generates and operates the management layer for multi-cloud deployments, tenant management, billing integration, and day-2 operations across every distribution channel. ## Common Distribution Challenges We Solve Teams across different roles face unique challenges when trying to distribute their software across multiple deployment models: - **Deploy Anywhere** - Go from App to Product by building your software distribution channels from On-Prem/Air-Gapped, BYOC, PaaS, SaaS, to Agentic SaaS in days. [Learn more](https://docs.omnistrate.com/usecases/overview/index.md) - **BYOC (Bring Your Own Cloud) Anywhere** - Deploy and manage your applications in customer cloud accounts and On-Prem with full control and isolation. [Learn more](https://docs.omnistrate.com/usecases/byoc/index.md) - **PaaS (Platform as a Service)** - Transform your software into a managed platform experience with automated provisioning, scaling, and multi-tenant management across any cloud infrastructure. [Learn more](https://docs.omnistrate.com/usecases/paas/index.md) - **SaaS (Software as a Service)** - Build fully managed cloud services with integrated billing, automated operations, and enterprise-grade security and compliance features. [Learn more](https://docs.omnistrate.com/usecases/saas/index.md) - **AaaS (Agent as a Service)** - Build infrastructure for autonomous AI systems with intelligent decision-making capabilities that adapt to customer needs, automate complex workflows, and optimize performance across distributed environments. [Learn more](https://docs.omnistrate.com/usecases/agent-aas/index.md) - **On-Premises & Air-Gapped** - Enable your customers to install applications in air-gapped environments. [Learn more](https://docs.omnistrate.com/usecases/air-gapped/index.md) - **Platform Extension** - Easily extend your existing platform stack to new deployment channels without replacing anything. [Learn more](https://docs.omnistrate.com/usecases/integrate-existing-stack/index.md) - **Open-Source to Open-Core** - Build your Open-Core Business Model for your Open-Source application without compromising your community-driven values. [Learn more](https://docs.omnistrate.com/usecases/open-source/index.md) - **Internal Delivery Platform** - Host your internal projects as managed services for internal teams with automated deployments and operational automation. [Learn more](https://docs.omnistrate.com/usecases/internal-saas/index.md) - **Build, List and Co-Sell with Cloud Marketplaces** - Unlock your Cloud Marketplace revenue pipeline by building a transactable and cloud-compliant Marketplace Offer, listing it on Cloud Marketplaces & Co-Sell with Cloud Providers in days. For specialized on-premises and air-gapped deployments, reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) ## Explore by Your Role Different teams have different needs when building and operating multi-cloud SaaS products. Find the resources most relevant to your role: ### For Platform Developers Build and integrate your applications with Omnistrate's platform. Learn how to integrate with our APIs. - [Tenancy SDK](https://docs.omnistrate.com/api/api-resources/index.md) - SDK for tenant management, identity, onboarding, controls - [Getting Started Guide](https://docs.omnistrate.com/getting-started/overview/index.md) - Get your first SaaS running in 5 minutes - [Omnistrate Build Guide](https://docs.omnistrate.com/build-guides/overview/index.md) - Comprehensive guide for building SaaS products - [API Documentation](https://docs.omnistrate.com/api/api-resources/index.md) - Complete API reference and examples - [Integrations](https://docs.omnistrate.com/build-guides/integrations/index.md) - Connect with your favorite tools and services - [Tenancy Guide](https://docs.omnistrate.com/tenant-management/overview/index.md) - Manage tenants, subscriptions and self-serve Customer Portal ### For Product & Business Teams Understand use cases, market opportunities, and business value of multi-cloud distribution. - [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) - Self-service customer management - [Use Cases](https://docs.omnistrate.com/usecases/overview/index.md) - Real-world implementation patterns - [Examples](https://github.com/omnistrate-community/examples) - Examples on services build with Omnistrate - [Customer Success Stories](https://blog.omnistrate.com/) - Learn from other successful implementations ### For DevOps & SRE Teams Automate deployments, manage infrastructure, and ensure reliable operations across all distribution channels. Module specific guides: - [Runtime Guide](https://docs.omnistrate.com/runtime-guides/overview/index.md) - manage capabilities that you can enable for your application - [DevOps Guide](https://docs.omnistrate.com/dev-ops-guides/overview/index.md) - capabilities to streamline your software development lifecycle - [Infra Guide](https://docs.omnistrate.com/infra-guides/overview/index.md) - manage your infrastructure from provisioning, updating, scaling, recovery, etc - [Operate Guide](https://docs.omnistrate.com/operate-guides/overview/index.md) - operate your distribution channels from operational perspective ### For FinOps & Business Operations Optimize costs, implement billing strategies, and manage financial operations across deployments. - [FinOps Guide](https://docs.omnistrate.com/fin-ops-guides/overview/index.md) - manage your distribution channels from cost perspective ### For SecOps & Governance Implement security controls, ensure compliance, and maintain governance across deployments. - [Governance guide](https://docs.omnistrate.com/governance-guides/overview/index.md) - enforce controls for security and compliance ## Why Choose Omnistrate? ### Distribution Channels Take Years to Build and an Army to Operate—Not Anymore Building a reliable control plane to manage your distribution channels is expert-level work that can consume your team's resources and often requires specialists. Different customers have different deployment needs, creating endless complexity across: - **Any Distribution Channel**: On-Prem/Air-Gapped, BYOC, PaaS, SaaS, Platform for AI Agents - **Any Tenancy Type**: Multi-Tenant, Dedicated VMs, Custom Networks, Dedicated Accounts - **Any Cloud, Any Region**: AWS, GCP, Azure, OCI, On-Premises across global regions - **Day-2 Operations**: Infrastructure management, high availability, fleet observability, DevOps automation, FinOps, and SecOps Omnistrate eliminates this complexity by generating a proven, enterprise-grade, private control plane that integrates seamlessly with your existing development stack, whether you're using Docker Compose, Helm Charts, Kubernetes Operators, Terraform, or OpenTofu. ### Get Started in Minutes, Not Months 1. **Bring Your Software** - Define your application with the artifacts you already use and get started in minutes 1. **Generate Your Control Plane** - Configure the APIs, Customer Portal, deployment workflows, tenancy, and operations model for your SaaS product 1. **Pay Only for What You Use** - Enjoy transparent, usage-based pricing with no hidden costs ## Connect with us Ready to transform your software into a multi-cloud SaaS? Here are several ways to get started and connect with our team: **See a Demo**\ Watch a live demonstration of how Omnistrate can transform your software into a multi-cloud SaaS in minutes. See a generated private control plane in action and learn how we can accelerate your go-to-market strategy. [Book a demo](https://calendly.com/omnistrate/meeting) or reach out to us at [demo@omnistrate.com](mailto:demo@omnistrate.com) **Build Your SaaS**\ Ready to see how your specific application would work on Omnistrate? Our team will help you build a proof-of-concept of your SaaS using our platform, showing you exactly how to deploy across multiple clouds and distribution channels.\ Reach out to us at [buildmysaas@omnistrate.com](mailto:buildmysaas@omnistrate.com) **Start Your Free Trial**\ Get hands-on experience with Omnistrate. Start generating your SaaS control plane today with our free trial—no credit card required. Begin transforming your software distribution in minutes. [Get started for free](https://omnistrate.cloud/signup) or reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) **Talk to Our Founders**\ Get direct access to our founding team for strategic guidance and technical insights. Schedule a personalized consultation to discuss your SaaS vision, explore partnership opportunities, and receive expert advice on building your control plane strategy. [Schedule a call](https://calendly.com/kamalwgupta/meeting) with one of our founders for 1-on-1 help building your first SaaS **Join Our Community**\ Connect with a vibrant community of SaaS builders, platform engineers, and industry experts. Share experiences, discover best practices, and get real-world insights from peers who are solving similar challenges in the SaaS space. Connect with us through the [Omnistrate support center](https://omnistrate.com/support) or LinkedIn [forum](https://www.linkedin.com/groups/13991058/) **Stay Updated**\ Stay informed about the latest developments in SaaS technology, platform engineering trends, and Omnistrate product updates. Access thought leadership content, industry insights, and company milestones. Read our [newsroom](https://omnistrate.com/press-release) and [blog](https://blog.omnistrate.com/), or follow us on [LinkedIn](https://www.linkedin.com/company/omnistrate/) # What is Omnistrate? ## Overview Modern software distribution has evolved far beyond simple code deployment to encompass complex multi-tenant SaaS offerings, hybrid cloud architectures, and diverse customer deployment preferences. Organizations struggle with fragmented toolchains that create operational silos, inconsistent customer experiences, and scalability bottlenecks. Whether customers demand public cloud SaaS, private cloud BYOC (Bring Your Own Cloud), or on-premises installations, delivering software across these varied environments while maintaining security, compliance, and cost efficiency requires a sophisticated orchestration layer. Without a unified control plane, software companies face exponential complexity in managing deployments, operations, and customer success across multiple infrastructure paradigms, ultimately limiting their ability to scale and compete effectively in today's diverse technology landscape. Omnistrate is an agentic control plane generator that enables software vendors to rapidly transform their software into multi-cloud SaaS products. With Omnistrate, you can generate a private control plane for your product and use it to build, distribute, and operate your software across multiple channels, including on-premises, BYOC (Bring Your Own Cloud), PaaS, SaaS, and AI agent deployments. A control plane is the management layer your teams and customers use to run the product lifecycle. It exposes APIs and user interfaces, provisions infrastructure, manages tenants and subscriptions, coordinates upgrades, tracks usage, applies access controls, and provides operational visibility. Omnistrate generates this layer from your product definition, then helps you run it under your brand, domain, and customer experience instead of exposing Omnistrate directly to your customers. Omnistrate enables you to build automated capabilities for complex tasks such as infrastructure provisioning, tenant management, billing integration, and day-2 operations. You can focus on your core product while using Omnistrate to generate and operate a private control plane that manages multi-cloud deployments, enterprise-grade security, compliance, and observability. ## Every Distribution Channel Needs a Control Plane Each distribution channel, whether SaaS, BYOC, Air-Gapped, or AI agents, requires its own deployment patterns, operational procedures, and customer interfaces. A control plane standardizes those requirements into a cohesive management layer. Omnistrate generates a custom, private control plane that provides consistent APIs, unified monitoring, and standardized operational workflows across all channels, enabling you to scale your distribution strategy without multiplying operational complexity. ## Deployment Hub - **Self Serve Deployments with Visibility & Controls** - **Multi-cloud SaaS & BYOC deployments** - **OnPrem Installer** - **Tenants, Subscriptions, & Billing** - **SaaS & Agentic experiences** Omnistrate generates a unified control plane that simplifies multi-environment software delivery. You can deploy applications across any infrastructure without requiring extensive DevOps expertise or custom integrations. The generated control plane lets both technical and non-technical users create self-service deployment experiences while maintaining enterprise-grade security and compliance. The platform automatically handles infrastructure provisioning, configuration management, and dependency resolution through intelligent orchestration engines. You can create parameterized deployments that customize per tenant while maintaining consistency. Real-time deployment tracking with rollback capabilities ensures zero-downtime updates, and integrated billing and subscription management automates the customer lifecycle from onboarding to revenue recognition. ## Operations Center - **Reliability, Security & Compliance** - **Ops Visibility & Control** - **Intelligent Upgrades** - **AI Feedback loop** Omnistrate enables you to build your own Operations Center that consolidates all operational aspects into a single, AI-powered dashboard. You can create monitoring and management capabilities for your entire fleet proactively, with the platform helping you predict issues before they impact customers. This approach reduces operational overhead while improving system reliability and customer satisfaction. The platform continuously analyzes system performance, user behavior, and infrastructure metrics to optimize operations automatically. Advanced anomaly detection identifies potential issues before they escalate, while intelligent upgrade systems ensure seamless updates with minimal disruption. Cost optimization engines automatically right-size resources based on actual usage patterns, and the AI feedback loop continuously improves recommendations and automation rules. ## How Omnistrate Helps You Generate Your Control Plane Building a control plane from scratch requires solving complex distributed systems challenges, managing multi-cloud infrastructure intricacies, and creating sophisticated automation layers that can handle the operational complexity of modern software distribution. Traditional approaches force engineering teams to become experts in infrastructure orchestration, tenant isolation, billing systems, and operational monitoring—diverting focus from core product innovation. The result is often fragmented solutions that struggle to scale, lack enterprise-grade reliability, and create maintenance burdens that grow exponentially with customer adoption. Omnistrate transforms this paradigm by generating the control plane from your SaaS product definition. Instead of assembling disparate tools and managing integration challenges, you define your product, deployment models, Resources, parameters, Tenancy Types, and operational policies. Omnistrate turns those inputs into private APIs, Customer Portal flows, lifecycle workflows, fleet operations, observability, billing hooks, and governance controls for enterprise-scale software distribution. ## Deployment Management Modern software vendors must deliver their products across a variety of deployment channels to meet their customer expectations. Some of your customers prefer the convenience of Hosted SaaS, some would prefer more control but yet fully managed through PaaS, others require even more control through Bring Your Own Cloud (BYOC), while regulated industries may demand on-premises or even air-gapped deployments. Omnistrate enables unified deployment management by letting you define your service once and distribute it seamlessly across all required channels. The same versioned service definition can be deployed as hosted SaaS, PaaS, in a customer’s cloud account (BYOC), on-premises, air-gapped—without duplicating engineering work. Each deployment inherits consistent lifecycle management, upgrades, and operational intelligence directly from the control plane. Each of the above deployment models can support multiple tenancy variants, depending on how isolation is enforced and how much control is delegated to the customer. For example: Hosted PaaS - Multi-tenancy across pods in the same cluster - Isolation at the VM/machine level - Isolation at the Kubernetes cluster level - Separation by network namespace or VPC - Full isolation per account BYOC (Bring Your Own Cloud) - Bring your own account (full isolation in the customer’s cloud account) - Bring your own VPC (running in the customer’s network boundary) - Bring your own Kubernetes cluster (control delegated to their cluster runtime) Across these models, tenancy is a spectrum—ranging from highly efficient multi-tenant pods to fully isolated accounts or physical environments. Omnistrate framework supports this entire spectrum out of the box, letting you define once how your service can be deployed, and letting your customers choose based on their needs. Alternatively, you can also build SaaS and completely abstract this to automatically adapt to the required tenancy level depending on customer demands, regulatory constraints, or SLA guarantees. In summary, Omnistrate turns distribution into a strategic growth lever—helping you reach more customers, in more environments, with no overhead. ## Infrastructure Management Modern software requires dynamic infrastructure that can adapt to varying workloads, scale across multiple cloud providers, and maintain consistency across diverse deployment environments. Traditional infrastructure-as-code tools like Terraform (or OpenTofu) provide static provisioning but lack the dynamic orchestration, dependency management, and operational intelligence needed for complex multi-tenant systems. Please see this [blog post](https://blog.omnistrate.com/posts/58) for more information on this topic. Omnistrate enables you to build infrastructure management capabilities that provision and manage your service infrastructure across any cloud providers, deployment models and tenancy types required for your service. The infrastructure required for your service can be defined with simple abstractions as part of your service definition, and all the required dependencies in the cloud provider (virtual machines, storage volumes, dns endpoints, load balancers, networking rules) are managed automatically through your custom control plane. Any change on the required infrastructure is considered in the service definition versioning and can be deployed safely using deployment pipelines following best practices. Any pre-requisite or cloud provided managed services that need to be configured before your service can run can also be specified in your control plane, allowing you to leverage any managed service from the cloud provider (ie RDS, DynamoDB, Lambda, SQS), while keeping a versioned control of the infra definition and supporting multiple deployment models for your service. ## Tenant Management Modern software distribution requires sophisticated tenant management capabilities and comprehensive customer lifecycle management to adapt to diverse customer requirements. Traditional approaches often result in rigid customer experiences that struggles to accommodate different customer requirements, deployment preferences, and operational needs, forcing compromises in operations and customer satisfaction. Omnistrate enables you to build comprehensive tenant management capabilities that handle the full customer lifecycle from onboarding to ongoing operations. The platform provides flexible tenancy models, automated deployment orchestration, and integrated billing systems that can adapt to your specific business requirements while maintaining enterprise-grade security and compliance and low operational cost. Learn more about Tenant Management [here](https://docs.omnistrate.com/tenant-management/overview/index.md) ## Operations Management Operating distributed software at scale requires sophisticated monitoring, automated recovery systems, and intelligent operational workflows that can handle complex failure scenarios while maintaining high availability. Traditional operational approaches often rely on reactive monitoring and manual intervention, creating operational bottlenecks that limit scalability and increase the risk of service disruptions. Omnistrate enables you to build comprehensive operational management capabilities that provide proactive monitoring, automated recovery, and intelligent fleet management. The platform continuously monitors your entire deployment fleet, automatically handles common failure scenarios, and provides operational teams with the visibility and tools needed to maintain enterprise-grade service levels. - [Monitoring with automated recovery](https://docs.omnistrate.com/operate-guides/monitoring/index.md): Handle failure scenarios to help achieve 99.99 SLA from simple machine failures to complex AZ failures, network partitions, process deadlatches etc - [SaaS Version management](https://docs.omnistrate.com/dev-ops-guides/upgrades/index.md): SaaS Product versioning encapsulating software, infrastructure and configuration to rollback or slowly rollout across tenants and their deployments - [Intelligent patching](https://docs.omnistrate.com/dev-ops-guides/upgrades/index.md): Manage your entire fleet from upgrading infrastructure, updating configuration to patching your software - [Inventory management](https://docs.omnistrate.com/operate-guides/overview/#inventory-management): Search control plane metadata and view associated information - [Operational visibility](https://docs.omnistrate.com/operate-guides/overview/#operational-visibility): Operational visibility into whats happening in your fleet - [Alert center](https://docs.omnistrate.com/operate-guides/overview/#alert-center): Configure notification types and channels to your needs - [Tenant management](https://docs.omnistrate.com/tenant-management/overview/index.md): Manage users, their subscription and deployments - [Observability integration](https://docs.omnistrate.com/build-guides/integrations/#observability-integrations): First-class integration with your favorite metrics and logging tools - [Alerting integration](https://docs.omnistrate.com/build-guides/integrations/#alarming): Enable real-time alerting with PagerDuty - [Automated fleet operations](https://docs.omnistrate.com/operate-guides/overview/index.md): Get L1 support spanning from automated operations - [Continuous deployment](https://docs.omnistrate.com/dev-ops-guides/pipelines/index.md): Setup Continuous Deployment - [CI integration](https://docs.omnistrate.com/dev-ops-guides/pipelines/#continuous-integration): Integration with Continuous Integration - [Access control for your teams](https://docs.omnistrate.com/governance-guides/omnistrate-rbac/index.md): Access control for internal teams for maintaining security, compliance, and operational efficiency within an organization - [Subscription management](https://docs.omnistrate.com/tenant-management/subscription-management/index.md): Simplify and automate the subscriptions for your service ## Application Management Most applications have dependencies on some open-source databases, queues, caching or other systems. In addition, most applications these days are not monolith and have multiple micro-services (or tools exposed to agents). All of these need to be packaged together. If you are already using Helm, Kustomize or Operator, you can bring your packaged artifact. However, some early-stage startups mayn't have the fully packaged product. Omnistrate automatically uses your docker compose to package your product effectively without you having to build your Helm, Kustomize or Operator. Here are the problems it addresses: - Defining dependency relationships for provisioning and upgrades - Defining placement of different resources to isolate respective components from each other - Defining the core Kubernetes resource manifests like ConfigMaps, Deployments, Secrets, Services, etc to map your final artifacts to Kubernetes concepts - Packaging your application and its dependencies (e.g., Redis, Postgres) as a single unit - Enabling versioning for its lifecycle management - Enable templatizing your deployments for your customers to customize deployments Moreover, modern applications require sophisticated runtime management capabilities that can handle dynamic scaling, data persistence, high availability, and complex networking requirements across diverse deployment environments. Traditional application management approaches often lack the intelligence and automation needed to optimize performance, ensure reliability, and maintain cost efficiency at scale. Omnistrate enables you to build comprehensive application management capabilities that provide intelligent scaling, automated backup and recovery, advanced networking, and performance optimization. The platform continuously monitors application performance and automatically adjusts resources to maintain optimal performance while minimizing costs. - [Autoscaling with custom metrics](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md): Auto-scaling with custom metrics - [Serverless scale down to zero](https://docs.omnistrate.com/runtime-guides/serverless/#scale-down-to-zero): Auto-scaling all the way down to zero - [Backup and Restore](https://docs.omnistrate.com/runtime-guides/pitr/index.md): Save data periodically and restore it when needed - [Custom tagging](https://docs.omnistrate.com/runtime-guides/custom-tagging/index.md): Organize and manage their cloud resources efficiently - [Load balancer support](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-load-balancer): Load balance across multiple nodes - [Stable Egress IP](https://docs.omnistrate.com/runtime-guides/overview/index.md): Provide stable egress IP for outbound traffic - [Multi-zone HA support](https://docs.omnistrate.com/runtime-guides/overview/index.md): Place nodes in different availability zones - [Reverse proxy with load balancing](https://docs.omnistrate.com/runtime-guides/overview/index.md): Enable https endpoint offload and add TLS termination proxy for secure communication over the internet - [Service discovery](https://docs.omnistrate.com/runtime-guides/overview/index.md): Find and communicate with each different components seamlessly - [Service account policies](https://docs.omnistrate.com/runtime-guides/overview/index.md): Securely talk to cloud-native services like Secrets Manager or MSK Connect ## Financial Operations and Monetization Financial operations and monetization for software products require sophisticated systems that can handle complex usage patterns, multiple pricing models, and diverse customer requirements across different deployment channels. Traditional billing approaches often lack the flexibility and automation needed to support modern SaaS business models, creating revenue leakage and operational overhead that scales poorly with customer growth. Omnistrate enables you to build comprehensive FinOps and monetization capabilities that provide accurate usage metering, flexible billing models, and seamless payment processing. The platform automatically tracks resource consumption, calculates charges based on your pricing models, and integrates with existing billing systems to streamline revenue operations. - [Metering](https://docs.omnistrate.com/fin-ops-guides/metering/index.md): Meter your per-customer usage - [Billing](https://docs.omnistrate.com/fin-ops-guides/billing/index.md): End-to-end billing including metering, aggregation, invoicing, payment, pricing plans and notifications - [Marketplace integration](https://docs.omnistrate.com/fin-ops-guides/marketplace/index.md): List your offerings to cloud marketplace - [External billing integration](https://docs.omnistrate.com/fin-ops-guides/billing/#byo-billing-provider-bring-your-own-billing-provider): Integrate with a billing system of your choice - [Cost Insights & Control](https://docs.omnistrate.com/fin-ops-guides/cost-insights/index.md) ## Enterprise-grade Security and Compliance Enterprise customers require rigorous security standards, comprehensive compliance frameworks, and advanced operational capabilities that meet the highest industry requirements. Traditional approaches to security and compliance often involve fragmented solutions that create gaps in coverage, increase operational complexity, and slow down product development cycles. Omnistrate enables you to build enterprise-grade security and compliance capabilities that provide comprehensive protection, automated compliance reporting, and advanced operational features. The platform includes built-in security controls, compliance frameworks, and enterprise integrations that accelerate your path to enterprise readiness. - [SOC2 for your control plane](https://docs.omnistrate.com/governance-guides/overview/index.md): Accelerate SOC2 for your SaaS Product - [Security questionnaire](https://docs.omnistrate.com/governance-guides/overview/index.md): Common questionnaire report for compliance like [AWS FTR checklist](https://aws.amazon.com/partners/foundational-technical-review/) - [Pen test report](https://docs.omnistrate.com/governance-guides/overview/index.md): Penetration testing - [SSO integration](https://docs.omnistrate.com/tenant-management/identity-providers/index.md): Integrate with common SSO providers for your customers to use SSO auth - [Endpoint aliases](https://docs.omnistrate.com/infra-guides/endpoint-aliases/index.md): Configure resource endpoint as alias - [IP whitelisting](https://docs.omnistrate.com/runtime-guides/overview/index.md): Whitelist incoming IPs for incoming traffic - [Advanced serverless](https://docs.omnistrate.com/runtime-guides/serverless/#advanced-serverless): Scale down to zero with full control - [Operational Status page](https://docs.omnistrate.com/build-guides/integrations/#operational-status): Operational status page to communicate to your users during outage at scale - [Cloud Insurance Integration](https://docs.omnistrate.com/build-guides/integrations/#cloud-insurance): Purchase cloud with unprecedented flexibility ## Reuse your existing investment Many organizations have already invested significant time and resources in building infrastructure automation, deployment pipelines, and operational tooling using technologies like Terraform, OpenTofu, Helm charts, or custom Kubernetes operators. The prospect of abandoning these investments to adopt a new platform can create resistance and delay adoption, even when the benefits are clear. Omnistrate recognizes and preserves your existing investments by providing seamless integration pathways for your current infrastructure and operational assets. Whether you've built sophisticated Terraform modules, developed comprehensive Helm charts, or created custom Kubernetes operators, you can bring these assets into your Omnistrate-powered control plane and continue building from where you are today. This integration approach means you don't lose your previous investment but instead accelerate your progress by combining your existing automation with Omnistrate's advanced orchestration, multi-cloud capabilities, and enterprise-grade operational features. You can start moving faster with lower effort while gradually expanding your control plane capabilities, allowing your team to focus on innovation rather than rebuilding foundational components from scratch. ## Start Simple, Scale with Confidence Building a comprehensive control plane doesn't require implementing every capability from day one. Software companies need the flexibility to start with their most critical requirements and expand their control plane capabilities as their business grows and customer needs evolve. Traditional platform approaches often force all-or-nothing decisions that can overwhelm teams and delay time-to-market. Omnistrate's modular design enables you to begin your control plane journey with the capabilities that matter most to your current stage—whether that's automating deployments, implementing basic tenant management, or establishing operational monitoring. As your customer base grows and requirements become more sophisticated, you can seamlessly add new capabilities like advanced billing systems, enterprise security features, or multi-cloud orchestration without rebuilding your foundation. This growth-oriented approach aligns with your business model through usage-based pricing that scales with your success. You only pay for the capabilities you use and the value you deliver to your customers, making Omnistrate a true partner in your journey from startup to enterprise-scale software company. As you expand into new markets, add distribution channels, or serve larger enterprise customers, Omnistrate grows with you, providing the advanced capabilities you need when you need them. To learn more on how Omnistrate allows you to do this, follow the link to [How it works](https://docs.omnistrate.com/what-you-can-do/index.md) # What You Can Do with Omnistrate ## Build distribution channels for any software Omnistrate solves the fundamental distribution channel problem that every software company faces: how to deliver your application across multiple deployment models, from on-premises and air-gapped environments to BYOC, PaaS, SaaS, and AI agent deployments, without rebuilding your infrastructure for each channel. Omnistrate is a control plane generator. It takes your existing application definitions and turns them into a comprehensive, private control plane capable of managing deployments, operations, and customer experiences across any infrastructure. Whether you've defined your application using container images, Docker Compose, Helm charts, Kubernetes operators, Kustomize configurations, or Terraform modules, Omnistrate builds upon these definitions to provide enterprise-grade distribution capabilities while maintaining your development team's control and flexibility. The generated control plane gives you the management layer that a SaaS product needs: APIs, customer portal, deployment workflows, tenant and subscription management, usage metering, access control, operational visibility, and fleet automation. Omnistrate helps you **build** your SaaS product distribution in minutes, **deploy** it in any model, **operate** it at scale, and **monetize** it wherever it runs. - **Build**: Start building your SaaS product distribution with Omnistrate. Omnistrate converts your application into a multi-channel SaaS product that can be deployed across any deployment model. More details, [here](#build-your-saas-product) - **Deploy**: Choose how to deploy your SaaS product across multiple channels, from self-serve Customer Portal flows to direct API integration, supporting any deployment model from on-premises to cloud. More details, [here](#deploy-your-saas-product) - **Operate**: Omnistrate provides an Operations Center that your DevOps, platform, and SecOps teams can use for day-2 operations, including fleet health, inventory, patching, visibility, and operational insights. You can also evolve your SaaS product by tuning the self-service customer experience, upgrading infrastructure, and testing changes in a sandbox environment. More details, [here](#operations-center) - **Monetize**: The generated control plane enables you to monetize your software across all distribution channels with integrated billing, usage metering, subscription management, and marketplace integrations, regardless of how your software is deployed. More details, [here](#finops-center) Let's see it in action with a practical example of what Omnistrate can do. ## Cloud Account Onboarding Before building your SaaS product, the first step is to onboard a cloud account. Omnistrate will deploy a control plane agent in your service account that is used to orchestrate your deployments. This onboarding process establishes the foundation for Omnistrate to manage your software deployments by creating the necessary infrastructure components within your cloud environment. The control plane agent acts as the orchestration layer that enables seamless deployment and management across Hosted SaaS, BYOC, and air-gapped deployment models. Importantly, your customers don't interact directly with Omnistrate, but with your control plane in your domain, ensuring a seamless branded experience. To learn how to onboard your cloud account you can follow [Account Onboarding Guideline](https://docs.omnistrate.com/getting-started/account-onboarding/index.md) ## Build your SaaS Product Your SaaS product is a collection of one or more Resources. Each Resource is a unit of functionality that can be run independently across any distribution channel. You can think of Resource as a microservice or a software component, with an optional Container Image, infrastructure or configuration attached. To learn more about Resource and what it entails, please see [this page](https://docs.omnistrate.com/build-guides/resource/index.md). Omnistrate acts as an abstraction layer that builds on top of your existing deployment definitions to generate your control plane while maintaining developer control. You can define your application and build your SaaS product distribution in multiple ways: - You can bring your own Code Base (GitHub repository), see [here](https://docs.omnistrate.com/getting-started/build-from-repo/index.md) - You can bring your own Docker Compose specification, see [here](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) - You can bring your Helm charts, see [here](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) - You can bring your Kubernetes Operators, see [here](https://docs.omnistrate.com/getting-started/build-from-operators/index.md) - You can bring your Kustomize configurations, see [here](https://docs.omnistrate.com/getting-started/build-from-kustomize/index.md) - You can bring your Terraform modules, see [here](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md) ## Docker Compose Example This example illustrates how Omnistrate works with Docker Compose, a tool for defining and running multi-container Docker applications using YAML configuration files (see [Docker Compose documentation](https://docs.docker.com/compose/) for more details). If you already have a Docker Compose specification, you can import it into Omnistrate. Omnistrate parses the specification and generates the equivalent control plane experience for your SaaS product. To set it up using Compose, please see [here](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) Note Docker Compose specification is just a way to quickly get started and is not a requirement to use Omnistrate. You are welcome to use the APIs/UI directly and explore other ways to build your SaaS product. Once imported, you can review your SaaS product on the Deploy portal. If you are happy with your SaaS product, you can deploy it to your customers. Otherwise, you can continue to [iterate through your SaaS product](https://docs.omnistrate.com/dev-ops-guides/upgrades/#how-to-make-the-changes), and Omnistrate will automatically manage all of your SaaS product versions. You can find more examples in our [Community Github repo](https://github.com/omnistrate-community/examples/tree/main/compose) Note We have provided specifications for some of the most common open source technologies across Docker Compose, Helm, Operators, Kustomize, and Terraform or OpenTofu. Feel free to leverage them or use them as examples to build your own. If you need help with any specification for one of the open-source softwares, please reach out by email to [support@omnistrate.com](mailto:support@omnistrate.com) We will use a hello world example using MySQL. Here is a sample Docker Compose specification for reference: ``` version: '3' services: db: image: mysql restart: always environment: MYSQL_ROOT_PASSWORD: your_password_here MYSQL_DATABASE: your_database_name_here MYSQL_USER: your_username_here MYSQL_PASSWORD: your_password_here volumes: - ./data:/var/lib/mysql ports: - '3306:3306' ``` ## Create New SaaS product from Command Line, UI or API Alternatively, you can also use our API or CTL to achieve the same if you prefer a programmatic way. - For CTL, please see [CTL Reference](https://docs.omnistrate.com/getting-started/getting-started-with-ctl/index.md) - For API, please see [API-Reference](https://api.omnistrate.cloud/docs/external/) You can also quickly start your project using available templates provided by the Omnistrate community. For example, the [container-based-template](https://github.com/omnistrate-community/container-based-template) repository offers a ready-to-use structure for containerized applications. These templates help you accelerate onboarding and standardize your SaaS product setup, allowing you to focus on your core functionality while leveraging best practices for deployment and distribution. If you don't have a Docker Compose specification for your project or do not prefer to use that, you can use our UI to build your SaaS product by directly importing your application. ## Command Line example The CTL is a command-line tool designed to help you build your SaaS product. It provides a convenient interface to interact with Omnistrate and execute operations on your SaaS products. If you already have a Compose file ready, CTL is the recommended method for building your SaaS product. You can install the Omnistrate CTL following these simple [steps](https://ctl.omnistrate.cloud/install/) ``` ./omnistrate-ctl _ __ __ ___ __ _ ___ (_)__ / /________ _/ /____ / _ \/ ' \/ _ \/ (_- **Cloud Accounts** in the [Omnistrate Console](https://omnistrate.cloud/system-settings/cloud-accounts). 1. Locate the cloud account you want to offboard. 1. Click **Delete** on the account configuration. 1. Wait for Omnistrate to finish platform cleanup. During this phase the account may remain in `Deleting` while Omnistrate removes deployment cells, networking, and other platform-managed resources. 1. Do **not** delete the onboarding artifacts from your cloud account while the account is still `Deleting`. Omnistrate may still need the bootstrap identities and trust configuration to complete cleanup. 1. Once Omnistrate marks the account as `Ready_to_offboard`, remove the onboarding artifacts from your cloud provider: #### AWS Delete the CloudFormation stack that was created during onboarding: 1. Open the [AWS CloudFormation Console](https://console.aws.amazon.com/cloudformation/). 1. Locate the Omnistrate onboarding stack. 1. Select the stack and choose **Delete**. 1. Wait for the stack deletion to complete. #### GCP Run the offboarding script using Cloud Shell: 1. Open [Google Cloud Shell](https://shell.cloud.google.com/). 1. Run the offboarding script provided during the onboarding process to remove the service account and IAM bindings. #### Azure Run the offboarding script using Azure Cloud Shell: 1. Open [Azure Cloud Shell](https://shell.azure.com/). 1. Run the offboarding script provided during the onboarding process to remove the service principal, role assignments, and related resources. #### Nebius Delete the service account keys from the Nebius console and remove the bindings from the account configuration on Omnistrate to revoke access. #### OCI Run the offboarding script provided during onboarding: 1. Open the [OCI Cloud Shell](https://cloud.oracle.com/) or a terminal with OCI CLI configured. 1. Run the offboarding script to remove the identity resources (users, groups, policies, and compartments) created during onboarding. ### If onboarding artifacts were deleted too early If the CloudFormation stack, GCP onboarding resources, Azure service principal, or OCI identity resources were deleted before the account reached `Ready_to_offboard`, cleanup can get stuck. Recommended recovery: 1. Restore or recreate the onboarding artifact if possible. 1. Retry the Omnistrate offboarding flow. 1. If the target account, project, or subscription was materially changed out of band, prefer a clean re-onboarding instead of trying to reuse stale bootstrap state. ### What offboarding cleans up When you delete the account configuration, Omnistrate automatically removes all deployment cells and their underlying infrastructure (control plane, networking, and system components). Note The cloud-provider-specific onboarding artifacts (CloudFormation stacks, GCP service accounts, Azure service principals, OCI identity resources) are not automatically deleted. Remove them manually only after Omnistrate reports that the account is ready to offboard. For any questions about the offboarding process, reach out to [support@omnistrate.com](mailto:support@omnistrate.com). # Building your Product using Compose ## Getting Started [Compose](https://docs.docker.com/compose/) is a powerful tool for defining and running multi-container applications and is used widely by many-many open source projects as a goto mechanism to deploy projects locally. With Compose, you can specify all the services, networks, volumes, and their relationships in a single, easy-to-read YAML file. To start using Compose, you'll need to define a omnistrate-compose.yaml file that describes your application's services and their configurations. Here is an Compose file to bring up Redis: ``` version: '3.9' services: redis: image: redis:alpine ports: - '6379:6379' ``` ## Validate your compose config Before you submit your compose spec, you may want to validate the compose specification by running it locally on your desktop: ``` docker-compose up ``` ## Extend your compose specification to build your distribution channel We have defined several compose tags to allow you to extend your compose specification to build your distribution channel, ex - SaaS. Here are some examples: - Configure your cloud provider account using `x-omnistrate-service-plan` tag - Make a Resource distributed with multiple replicas using `x-omnistrate-compute` - Add a reverse proxy to HTTP-based services or other capabilities using `x-omnistrate-capabilities` - Customize visibility, logging, metering of a service using: `x-customer-integrations` - Customize API parameters for launching the service using `x-omnistrate-api-params` - Inject custom code at different phases of your SaaS using `x-omnistrate-actionhooks` - Customize the compute / network / storage parameters of a service via `x-omnistrate-compute` and other corresponding tags For more details on all the compose tags, please read more about [Omnistrate Compose Extensions](https://docs.omnistrate.com/build-guides/compose-spec/index.md) ### Use AI to easily extend your compose specification ``` Extend my compose specification to build my PaaS service with dedicated tenancy and basic observability ``` ### Example compose config ## Build your Product using your compose specification Create a Compose specification file and build your service: ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Your SaaS Product Name" --description "Your service description" ``` ## More Real-World Examples List of examples demonstrating how to build various types of applications using Omnistrate. Each example provides step-by-step guidance, complete configuration files, and best practices for different use cases. - [Vector database example using PostgreSQL](https://github.com/omnistrate-community/examples/blob/main/docs/dbaas/README.md) - [Serverless master-replica database example using MySQL](https://github.com/omnistrate-community/examples/blob/main/docs/mysql-master-replica-serverless/README.md) You can access the full examples index [here](https://github.com/omnistrate-community/examples/blob/main/docs/README.md). ## Dive Deeper Now that you have a working Compose-based SaaS Product, explore the build guides to add production-grade capabilities: - [Compose deployment strategy](https://docs.omnistrate.com/build-guides/compose-overview/index.md) — Browse the Compose-specific strategy and specification reference - [Compose Troubleshooting](https://docs.omnistrate.com/build-guides/compose-troubleshooting/index.md) — Diagnose failed Compose deployments and workflows - [Compose Specification Reference](https://docs.omnistrate.com/build-guides/compose-spec/index.md) — Complete reference for all Omnistrate-specific Compose extensions (`x-omnistrate-*` tags) - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Let your customers configure deployments with custom parameters - [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) — Use dynamic system-generated values in your Compose spec - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) — Configure hosted, BYOC, or air-gapped deployment models - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) — Choose between shared, dedicated, or hybrid tenancy - [Action Hooks](https://docs.omnistrate.com/build-guides/actionhooks/index.md) — Run custom logic during provisioning, scaling, or upgrades - [Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md) — Set up dependencies between resources for multi-component applications ## Additional Resources - [Compose Documentation](https://docs.docker.com/compose/) — The official documentation for Docker Compose - [Compose File Reference](https://docs.docker.com/compose/compose-file/) — Detailed guide to the Compose file format - [Docker Hub](https://hub.docker.com/) — Explore container images for use with your Compose specification # Building your SaaS Product using Helm charts ## Getting Started Omnistrate supports deploying Helm Charts as part of your service topology. Helm is a package manager for Kubernetes that allows you to define, install, and manage Kubernetes applications. This gives you more control over the Kubernetes manifests that are deployed as part of your service topology. In addition, you can bring your existing Helm charts and deploy them on Omnistrate without any modifications. As part of the deployment, Omnistrate takes care of the following: - Deploying a VPC / Subnets in the chosen region and chosen account (customer's or yours) - Deploying a Kubernetes cluster in the chosen region - Deploying NLBs w/ Nginx Ingress Controllers - Deploying a Kubernetes Dashboard for you to monitor your deployments - Deploying a Route53 Hosted Zone for your workload endpoints that you can configure through Kubernetes Service annotations - Deploying an IAM role / Google Service Account for your workload to invoke Cloud Provider APIs / Services like S3 - Deploying a Kubernetes Role / RoleBinding for your workload to manage Kubernetes resources within the namespace of the deployment - Configuring ACME TLS certificates that are auto-rotated - Deploying your Helm charts with any customer specific configurations Omnistrate fully supports these deployments as long as they are in a remote artifactory accessible to your deployment Kubernetes environment. If your product depends on cluster-wide prerequisites that should exist once per Kubernetes cluster, such as Prometheus, ExternalDNS, CSI drivers, or shared controllers, install those at the deployment-cell layer through [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) rather than as ad hoc per-instance post-install steps. ## Anatomy of a Helm Chart Helm chart registration on Omnistrate requires the following: - **Chart Name**: The name of the Helm chart. - **Chart Version**: The version of the Helm chart. - **Chart Repository**: The Helm chart repository URL. - **Chart Values**: The values file for the Helm chart if you want to override the default values. - **Endpoint Configuration**: The endpoint configuration for the Helm chart if you want to expose the service connectivity details to your customers. ## Registering a Helm Chart Helm charts are managed through a specification file that defines your overall service topology on Omnistrate. A complete description of the Plan specification can be found on [Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) Here is an example of a SaaS Product that deploys Redis Clusters using a Helm chart. ``` name: Redis Server # Plan Name deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: Redis Cluster network: ports: - 6379 endpointConfiguration: cluster: host: "$sys.network.externalClusterEndpoint" ports: - 6379 primary: true networkingType: PUBLIC admin: host: admin-{{ $sys.network.internalClusterEndpoint }} ports: - 8888 primary: false networkingType: PRIVATE helmChartConfiguration: runtimeConfiguration: disableHooks: false wait: true waitForJobs: true recreate: false resetThenReuseValues: false resetValues: true reuseValues: false skipCRDs: false upgradeCRDs: true timeoutNanos: 180000000000 chartName: redis chartVersion: 24.1.0 chartRepoName: bitnamicharts chartRepoURL: oci://registry-1.docker.io/bitnamicharts chartValues: global: security: allowInsecureImages: true image: registry: docker.io repository: bitnamisecure/redis tag: latest master: persistence: enabled: false resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi replica: persistence: enabled: false replicaCount: 1 resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi ``` ## Helm Chart specification breakdown A quick breakdown of the above specification: - **name**: Omnistrate allows you to define "plans" for your services that your customers subscribe to for different pricing / Tenancy Types (e.g. "Basic", "Pro", "Enterprise"). This is the name of the Plan. - **deployment**: This is the deployment configuration for the service. In this case, it is a Hosted SaaS deployment on AWS which hosts all your customer's Redis clusters in the specified AWS account. - **services**: This is the list of services that are part of the Plan. In this case, we have a single service called "Redis Cluster". - **network**: This is the network configuration for the service. In this case, we are exposing port 6379 for the Redis Cluster. Omnistrate fully manages your infrastructure and this specification defines what firewall / security group rules to apply. - **helmChartConfiguration**: This is the Helm chart configuration for the service. In this case, we are deploying the Redis Helm chart version 24.1.0 from the Bitnami Helm chart repository. We also override the default values for the Redis Helm chart to disable persistence and set resource limits for the master and replica pods. - **endpointConfiguration**: This is the endpoint configuration for the service. In this case, we are exposing the Redis Cluster on a public endpoint and the admin interface on a private endpoint. These are parameters that are exposed to your customers to connect to the service. Depending on the configuration of the Kubernetes service / ingress in your Helm chart, you can expose these endpoints to your customers. - **runtimeConfiguration**: This is the runtime configuration for the Helm chart. This is how Omnistrate manages the deployment of the Helm chart. With this, you can customize the runtime behavior of the Helm chart client. For more information, see: https://helm.sh/docs/helm/helm_install/#options Info You can use system parameters to customize Helm Chart values. A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). Info For advanced configuration scenarios, consider using [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) which provide conditional deployment customization across different environments and cloud providers. ## Pull Helm Charts from Private Amazon ECR For AWS Helm resources, `chartRepoURL` can point to a private Amazon ECR OCI repository. Do not add a static `authProvider` block for this path; Omnistrate resolves short-lived ECR credentials at deployment time through AWS IAM. ``` name: Redis Server deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: "arn:aws:iam:::role/omnistrate-bootstrap-role" services: - name: Redis Cluster helmChartConfiguration: chartName: redis chartVersion: 22.0.7 chartRepoName: private-ecr-charts chartRepoURL: oci://.dkr.ecr..amazonaws.com/ ``` Use the same `helmChartConfiguration` shape for `hostedDeployment` and `byoaDeployment`; choose the deployment block that matches your Plan. For Hosted SaaS deployments, the private ECR chart repository must be in the Service Provider / ISV AWS account configured on `hostedDeployment`. For BYOC deployments, the private ECR chart repository must be in the Service Provider / ISV AWS account used by the BYOC control plane, not the end customer workload account. For AWS account setup with CloudFormation, keep `EnableECRHelmChartPull=true` on the Service Provider / ISV AWS account that owns the private ECR repository. This is the Hosted SaaS deployment account for `hostedDeployment` Plans and the BYOC control plane account for `byoaDeployment` Plans. If the account was onboarded before this parameter was enabled, update the generated CloudFormation stack and keep this parameter set to `true`. ## Using a Local Helm Chart Artifact with CTL Use `artifactRelativePath` when the Helm chart is in your local workspace and you want `omnistrate-ctl build` to upload it as part of the Plan build. Run `omnistrate-ctl build` from the workspace root that contains the relative artifact path. CTL resolves `artifactRelativePath` relative to the current working directory, creates or reads a `.tar.gz` / `.tgz` artifact, uploads it, and then builds the Plan. Example workspace: ``` my-service/ spec.yaml charts/ redis/ Chart.yaml values.yaml templates/ ``` Example local-artifact spec: ``` name: Redis Server deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: "arn:aws:iam:::role/omnistrate-bootstrap-role" services: - name: Redis Cluster network: ports: - 6379 endpointConfiguration: cluster: host: "$sys.network.externalClusterEndpoint" ports: - 6379 primary: true networkingType: PUBLIC helmChartConfiguration: artifactRelativePath: charts/redis chartValues: replica: replicaCount: 1 ``` You can also point `artifactRelativePath` at an already packaged chart: ``` helm package charts/redis --destination local-artifacts ``` ``` helmChartConfiguration: artifactRelativePath: local-artifacts/redis-0.1.0.tgz chartValues: replica: replicaCount: 1 ``` Build the Plan with CTL: ``` cd my-service omnistrate-ctl build \ --spec-type ServicePlanSpec \ --file spec.yaml \ --product-name "RedisHelm" \ --release-as-preferred ``` For local chart artifacts, Omnistrate reads chart metadata from `Chart.yaml`; you do not need to set `chartRepoName` or `chartRepoURL`. Use one artifact source per resource: do not mix `artifactRelativePath` with `chartRepoName` / `chartRepoURL` for the same chart. ## Send Helm Resource Logs to CloudWatch Logs For Helm resources, choose the Cloud Native provider for internal logs on the Plan to send logs to CloudWatch Logs. The destination and credentials depend on the deployment model: - **Provider-hosted Helm:** Supported only when the instance is deployed on AWS. Logs are sent to CloudWatch Logs in the SaaS Provider's AWS deployment account using the deployment cell's native AWS credentials. - **BYOC Helm, including BYOC On-Premise:** Logs are sent to CloudWatch Logs in the SaaS Provider's AWS account, regardless of the dataplane cloud. Omnistrate supplies scoped, temporary AWS credentials to the collector. The following example configures a GCP BYOC dataplane and the SaaS Provider AWS account configured on `byoaDeployment`: ``` name: Redis Server deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: "arn:aws:iam:::role/omnistrate-bootstrap-role" features: INTERNAL: logs: provider: native services: - name: Redis Cluster compute: instanceTypes: - name: e2-small cloudProvider: gcp helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnamicharts chartRepoURL: oci://registry-1.docker.io/bitnamicharts ``` In the Omnistrate Console, select **Cloud Native** for internal logs. In the Plan specification, the same setting is `features.INTERNAL.logs.provider: native`. Use the same `features` block for `hostedDeployment` and `byoaDeployment`; choose the deployment block that matches your Plan. A provider-hosted Helm instance with native internal logs must be deployed on AWS. The CloudFormation flags for private Helm charts and native logs are independent. Keep `EnableECRHelmChartPull=true` when the Plan pulls a Helm OCI chart from the SaaS Provider's private ECR repository, as described in [Pull Helm Charts from Private Amazon ECR](#pull-helm-charts-from-private-amazon-ecr). Keep `ExpandBoundaryForAWSServiceIntegrations=true` when the Plan enables native internal logs. A Plan that uses both capabilities requires both flags. For provider-hosted AWS deployments, keep `ExpandBoundaryForAWSServiceIntegrations=true` in the SaaS Provider deployment account's generated CloudFormation stack. This grants the deployment cell's EKS node role permission to create CloudWatch log groups and streams, set retention, describe CloudWatch Logs resources, and put log events. If the account was onboarded before this parameter was enabled, update the generated CloudFormation stack and keep this parameter set to `true`. For BYOC deployments, keep `ExpandBoundaryForAWSServiceIntegrations=true` in the SaaS Provider AWS account's generated CloudFormation stack so Omnistrate can provide scoped, temporary AWS credentials. If the account was onboarded before this parameter was enabled, update the generated CloudFormation stack and keep this parameter set to `true`. The credentials allow writes only to the instance's CloudWatch log groups and stream prefix in the SaaS Provider's AWS account. The customer workload account does not need native logging permissions for this feature. The BYOC dataplane must have an HTTPS network path to the regional AWS CloudWatch Logs endpoint. It can use public egress or private, routed connectivity to an AWS CloudWatch Logs interface VPC endpoint. Cross-cloud and on-premise private paths require appropriate routing and DNS, such as connectivity through a VPN or AWS Direct Connect. Native logs are not supported from dataplanes that cannot reach CloudWatch Logs. When an instance is deployed, Omnistrate creates a per-instance OpenTelemetry collector for the Helm resource. For provider-hosted AWS deployments, the collector writes to CloudWatch Logs using the deployment cell's native AWS credentials. For BYOC deployments, it writes to CloudWatch Logs in the SaaS Provider's AWS account using scoped, temporary credentials. CloudWatch log groups are derived from the SaaS Product and Plan names, and log streams are prefixed with the instance ID. To switch back to Omnistrate-managed logs, remove `provider: native` and leave the internal logs feature enabled: ``` features: INTERNAL: logs: {} ``` ## Override Helm release name and namespace per deployment By default, Omnistrate chooses the Helm release name and Kubernetes namespace for each deployment. If your chart or operating model requires customer-controlled names, define API parameters and reference them from `helmChartConfiguration`. Use `releaseName` and `namespace` with `$var.` when an API parameter should directly provide the Helm release name or namespace: ``` services: - name: Redis Cluster helmChartConfiguration: chartName: redis chartVersion: 19.6.2 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami namespace: $var.helmNamespace releaseName: $var.helmReleaseName chartValues: replica: replicaCount: 1 apiParameters: - key: helmNamespace description: Helm namespace override name: Kubernetes Namespace type: String modifiable: true required: false export: true defaultValue: redis-namespace - key: helmReleaseName description: Helm release name override name: Helm Release Name type: String modifiable: true required: false export: true defaultValue: redis-release ``` When customers create an instance through the generated API or Customer Portal, they provide `helmReleaseName` and `helmNamespace`. Omnistrate renders the existing `releaseName` and `namespace` fields before installing, upgrading, or uninstalling the Helm chart. You can also set `releaseName` and `namespace` directly when you need static values or interpolation with API parameters and system parameters: ``` services: - name: Redis Cluster helmChartConfiguration: chartName: redis chartVersion: 19.6.2 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami namespace: "override-{{ $var.helmNamespace }}/{{ $sys.id }}" releaseName: "override-{{ $var.helmReleaseName }}/{{ $sys.id }}" apiParameters: - key: helmNamespace description: Helm namespace override name: Helm Namespace type: String modifiable: true required: false export: true defaultValue: redis-namespace - key: helmReleaseName description: Helm release name override name: Helm Release Name type: String modifiable: true required: false export: true defaultValue: redis-release ``` References are validated during service plan registration. Any `$var.` reference must match an API parameter on the service, and any `$sys.` reference must be a supported system parameter. The rendered release name and namespace are passed to Helm at deployment time. Info For more detailed information on the **pricing**, **metering**, **billingProviders** configuration, please see [End-to-End Billing](https://docs.omnistrate.com/fin-ops-guides/billing/index.md) and [Usage Metering](https://docs.omnistrate.com/fin-ops-guides/metering/index.md). And that's all it takes to setup: - A self-service REST API for your customers to deploy Redis Clusters. - A Hosted SaaS deployment on AWS to host all your customer's Redis Clusters. ## Registering a Helm Chart specification You can register this spec using our CLI: ``` omnistrate-ctl build -f spec.yaml --name 'RedisHelm' --release-as-preferred --spec-type ServicePlanSpec # Example output shown below ✓ Successfully built service Check the Plan result at: https://omnistrate.cloud/product-tier?serviceId=s-xxxxx&environmentId=se-xxxxx Access your SaaS Product at: https://saasportal.instance-xxxxx.hc-xxxxx.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-xxxxx&environmentId=se-xxxxx ``` ## Deploying a Redis Cluster through your dedicated Customer Portal Once you have registered the Helm chart, you can deploy a Redis Cluster using the API / UI we generate for your customers. We also setup a Dev environment for you to test your deployment before going live. In the above output, you can see the URL to access the Customer Portal for your Dev environment where you can deploy Redis Clusters just like your customers would. Sign-in using your existing Omnistrate credentials (if you signed up using SSO, you can go to your [profile](https://omnistrate.cloud/settings?view=Profile) in your Omnistrate portal and set a new password). You will be presented with a screen that defines the default plan you setup in the specification and a dashboard for you to manage your deployments. Let's deploy a cluster in AWS in us-east-1. Once the deployment is complete, you will see the status of the deployment in the dashboard. **Let's take a peek at our kubernetes cluster that was automatically setup by Omnistrate in the chosen region.** This gives you a quick overview of the process to design your Helm chart-based SaaS Product on Omnistrate. For more customizations please see [Helm Chart Customizations](https://docs.omnistrate.com/build-guides/helm-charts-customize/index.md). ## More Real-World Examples - [Helm-based PostgreSQL deployment](https://github.com/omnistrate-community/postgres-paas/tree/main/helm) - [Agent Runtime SaaS with Helm](https://github.com/omnistrate-community/examples/tree/main/docs/terraform-helm#helm-chart-service-agent-runtime) ## Status Monitoring and Endpoints Omnistrate monitors the health and status of your Helm chart deployments and exposes endpoint information to your customers. ### How status monitoring works For Helm chart deployments, Omnistrate tracks the status of your Kubernetes resources (Pods, Services, StatefulSets) to determine whether your deployment is healthy and ready to serve traffic. This includes: - **Pod readiness**: Monitoring pod status and readiness probes to detect when your application is available. - **Endpoint resolution**: Automatically resolving and exposing service endpoints (Load Balancer IPs, DNS names) to your customers. - **Health transitions**: Detecting transitions between healthy, degraded, and unhealthy states, and reflecting these in the Operations Center. ### Configuring endpoints Use `endpointConfiguration` in your Plan specification to define which endpoints are exposed to your customers: ``` services: - name: Redis Cluster endpointConfiguration: cluster: host: "$sys.network.externalClusterEndpoint" ports: - 6379 primary: true networkingType: PUBLIC ``` Omnistrate resolves `$sys.network.externalClusterEndpoint` once the Kubernetes Service has an external IP or DNS name assigned, and surfaces this to your customers through the Customer Portal and API. Important behavior to keep in mind: - `endpointConfiguration` describes the connectivity details Omnistrate should surface to customers. It does **not** create the Kubernetes `Service`, `Ingress`, load balancer, or DNS record by itself. - `networkingType: PUBLIC` means the endpoint is intended to be reachable over a public network path. `networkingType: PRIVATE` means the endpoint is intended for private connectivity workflows. - If you expose multiple services, give each one a distinct hostname that matches the DNS records your chart creates. - Endpoint information appears only after the underlying Kubernetes resources are provisioned and the host/IP is actually resolvable. When your chart uses `external-dns` annotations or cloud load balancer annotations, keep those aligned with the `host` values you publish in `endpointConfiguration`. ### Monitoring in the Operations Center You can monitor the status of all your Helm-based deployments from the **Operations Center**: - **Instance status**: View real-time health of each deployment instance. - **Workflow tracking**: Monitor provisioning, upgrade, and scaling workflows with detailed step-by-step progress. - **Endpoint availability**: Verify that endpoints are resolved and accessible. Tip If an instance remains in a provisioning state, check the workflow details in the Operations Center to identify which step is pending. Common causes include pods failing readiness probes or Kubernetes Services (or cloud load balancers) waiting for external IP or DNS assignment. ## Dive Deeper Now that you have a working Helm-based SaaS Product, explore the build guides to customize and extend your deployment: - [Helm deployment strategy](https://docs.omnistrate.com/build-guides/helm-charts-overview/index.md) — Browse the Helm-specific strategy and guide set - [Helm Troubleshooting](https://docs.omnistrate.com/build-guides/helm-charts-troubleshooting/index.md) — Diagnose failed Helm deployments and workflows - [Helm Chart Customization](https://docs.omnistrate.com/build-guides/helm-charts-customize/index.md) — Override chart values, configure affinity rules, and use automatic affinity injection - [Helm Chart Topology](https://docs.omnistrate.com/build-guides/helm-charts-topology/index.md) — Design multi-component service topologies with Helm - [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) — Apply conditional, environment-specific Helm values - [Helm Runtime Configuration](https://docs.omnistrate.com/build-guides/helm-charts-runtime-configuration/index.md) — Fine-tune Helm install and upgrade behavior - [Helm and Terraform](https://docs.omnistrate.com/build-guides/helm-charts-terraform/index.md) — Combine Helm deployments with Terraform-managed infrastructure - [Plan Specification Reference](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) — Complete reference for the Plan spec format used by Helm, Operators, and Terraform deployments - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Expose customer-configurable parameters through your SaaS APIs - [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) — Inject dynamic platform values into your Helm chart values # Building your SaaS Product using Kustomize ## Getting started Omnistrate supports deploying Kustomize-based configurations as part of your service topology. Kustomize is a configuration management tool for Kubernetes that allows you to customize raw, template-free YAML files for Kubernetes deployments. This provides flexibility when handling multiple environments and customizing resources. As part of the deployment, Omnistrate manages the following: - Deploying a VPC / Subnets in the chosen region and chosen account (customer's or yours) - Deploying a Kubernetes cluster in the chosen region - Deploying NLBs w/ Nginx Ingress Controllers - Deploying a Kubernetes Dashboard for you to monitor your deployments - Deploying a Route53 Hosted Zone for your workload endpoints that you can configure through Kubernetes Service annotations - Deploying an IAM role / Google Service Account for your workload to invoke Cloud Provider APIs / Services like S3 - Deploying a Kubernetes Role / RoleBinding for your workload to manage Kubernetes resources within the namespace of the deployment - Configuring ACME TLS certificates that are auto-rotated - Deploying your Helm charts with any customer specific configurations Omnistrate fully supports these deployments as long as they are in a remote repository accessible to your deployment Kubernetes environment, rendering and deploying Kustomize templates with any customer or environment specific configurations. ## Integrating Kustomize on your Plan Before deploying your service, you should prepare a Kustomize stack, which enables various customizations for each deployment. Here is an example of a Kustomize stack designed for a specific deployment: kustomization.yaml ``` resources: - pg.yaml - pgpv.yaml - pgpvc.yaml namespace: "{{ $sys.id }}" configMapGenerator: - name: pg-config literals: - defaultPassword=admin - pgDefaultUsername={{ $var.username }} - pgDefaultPassword={{ $var.password }} ``` pg.yaml ``` apiVersion: apps/v1 kind: Deployment metadata: name: postgres-deployment spec: replicas: 1 selector: matchLabels: app: postgres template: metadata: labels: app: postgres spec: containers: - name: postgres image: postgres:13 ports: - containerPort: 5432 env: - name: POSTGRES_DB value: exampledb - name: POSTGRES_USER valueFrom: configMapKeyRef: name: pg-config key: pgDefaultUsername optional: false - name: POSTGRES_PASSWORD valueFrom: configMapKeyRef: name: pg-config key: pgDefaultPassword optional: false volumeMounts: - mountPath: /var/lib/postgresql/data name: postgres-storage subPath: postgres affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/resource operator: In values: - '{{ $sys.deployment.resourceID }}' volumes: - name: postgres-storage persistentVolumeClaim: claimName: "{{ $sys.id }}-pvc" ``` pgpv.yaml ``` apiVersion: v1 kind: PersistentVolume metadata: name: "{{ $sys.id }}-pv" spec: capacity: storage: 5Gi accessModes: - ReadWriteOnce persistentVolumeReclaimPolicy: Delete storageClassName: gp2 hostPath: path: /mnt/data/postgres ``` pgpvc.yaml ``` apiVersion: v1 kind: PersistentVolumeClaim metadata: name: "{{ $sys.id }}-pvc" spec: accessModes: - ReadWriteOnce storageClassName: omnistrate-platform-default resources: requests: storage: 5Gi ``` Note For AWS EKS workloads that need EBS volumes encrypted with a customer-managed AWS KMS key, create a custom storage class and reference it with `storageClassName` in your PVC. See [AWS Storage Classes](https://docs.omnistrate.com/infra-guides/storage-classes/#aws-storage-classes). Kustomize templates need to be available on a Github repository. One repository can contain multiple Kustomize stacks, and can be referenced using a specific git reference (tag or branch) and a folder patch within that repository. Omnistrate allows you to configure a reference to a Git repository and path where the Kustomize stack is stored. When selecting a path for your repository Omnistrate expects the entire Kustomize definition to be under that folder structure. Within the Kustomize templates, the Omnistrate platform provides system parameters that can be used to reference to information about the current cluster, for instance: - `$sys.deploymentCell.publicSubnetIDs[i].id` - `$sys.deploymentCell.privateSubnetIDs[i].id` - `$sys.deploymentCell.region` - `$sys.deploymentCell.cloudProviderNetworkID` - `$sys.deploymentCell.cidrRange` You can inject these values into your Kustomize templates, and these values will be dynamically rendered during deployment. Info You can use system parameters to customize Kustomize templates. A detailed list of system parameters be found on [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ## Registering a Service using a Kustomize Stack Kustomize stacks are managed through a specification file that defines your overall service topology on Omnistrate. A complete description of the Plan specification can be found on [Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) Here is an example of using Kustomize to configure a SaaS Product: ``` name: Kustomize deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' services: - name: kustomizeRoot compute: instanceTypes: - name: t4g.small cloudProvider: aws - name: e2-medium cloudProvider: gcp - name: Standard_B2als_v2 cloudProvider: azure network: ports: - 5432 kustomizeConfiguration: kustomizePath: /artifacts/kustomize gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/omnistrate-community/examples.git apiParameters: - key: username description: Username name: Username type: String modifiable: true required: false export: true defaultValue: username - key: password description: Default DB Password name: Password type: String modifiable: false required: false export: false defaultValue: postgres ``` Info For more detailed information on the **deployment**, **pricing**, **metering**, **billingProviders**, **maxNumberOfInstancesAllowed** configuration, please see [here](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-service-plan). You can register this spec using our CLI: ``` omnistrate-ctl build -f spec.yaml --name 'Kustomize' --release-as-preferred --spec-type ServicePlanSpec ``` ## Dive Deeper Now that you have a working Kustomize-based SaaS Product, explore the build guides to add advanced capabilities: - [Plan Specification Reference](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) — Complete reference for the Plan spec format, including the Kustomize configuration schema - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Expose customer-configurable parameters through your SaaS APIs - [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) — Use dynamic system-generated values in your Kustomize templates - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) — Configure hosted, BYOC, or air-gapped deployment models - [Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md) — Set up dependencies between resources for multi-component applications - [Action Hooks](https://docs.omnistrate.com/build-guides/actionhooks/index.md) — Run custom logic during provisioning, scaling, or upgrades # Build from Kubernetes Operators Use this guide when you already have a Kubernetes Operator, or you want to build your product lifecycle around an operator-managed custom resource. Omnistrate does not replace your operator. Your operator continues to reconcile the application. Omnistrate provides the surrounding control plane: service plan versioning, tenant and subscription management, Customer Portal, generated APIs, cloud account onboarding, deployment cells, networking, observability, backups, restores, upgrades, and fleet operations. For the complete current example, start from the [operator spec template](https://github.com/omnistrate-community/operator-spec-template). It defines a CloudNativePG PostgreSQL service plan using operator Helm dependencies, `systemWorkflows`, backup configuration, snapshot metadata, and restore workflows. ## Start with the Omnistrate Operator Skill Use the `omnistrate-operator` skill when an AI assistant is helping you onboard a Kubernetes operator. The skill guides the assistant through operator discovery, Plan specification (`ServicePlanSpec`) authoring, `systemWorkflows`, deployment, and lifecycle validation. Before you begin: - Install `omnistrate-ctl` using the [installation guide](https://docs.omnistrate.com/getting-started/installing-ctl/index.md). - Authenticate with `omnistrate-ctl login`. - Use an AI assistant that supports Agent Skills. Install the skill with the skills CLI: ``` npx skills add omnistrate-oss/agent-instructions \ -s omnistrate-operator ``` For Claude Code, Codex, GitHub Copilot, and other supported assistants, see [Installing the Skills](https://github.com/omnistrate-oss/agent-instructions#installing-the-skills) in the Agent Instructions repository. The skill uses `omnistrate-ctl` by default. If you want the assistant to call Omnistrate tools directly, configure the optional [Omnistrate MCP server](https://docs.omnistrate.com/getting-started/mcp-server/index.md). The skill provides the operator onboarding workflow; MCP provides an agent connection to Omnistrate tools. MCP is not required when the assistant can run the CLI. ### Start an onboarding session Give the assistant the following information along with the `omnistrate-operator` skill: ``` Use the omnistrate-operator skill to onboard my Kubernetes operator to Omnistrate. Operator: Operator Helm chart: CRD group/version/kind: Example custom resource: Deployment model: Hosted SaaS / BYOC Anywhere Cloud and region: Required lifecycle operations: create, modify, stop, start, backup, restore, delete ``` Also provide the operator scope and watch namespace, CRD status conditions, operator-native stop/start mechanism, service names and ports, and cloud-account details. These facts let the assistant author a ServicePlanSpec against the operator's actual behavior instead of guessing its schema. The skill will: 1. Determine the operator's scope and select the correct CRD and controller installation pattern. 1. Create a minimal ServicePlanSpec with `systemWorkflows` for the first supported lifecycle operations. 1. Build and deploy the first instance, then debug workflow and custom-resource reconciliation failures. 1. Add lifecycle operations, networking, endpoints, backups, and production settings incrementally. 1. Validate the live custom resource and operator status before you harden the service for production. ## When to Use This Path Choose the operator path when: - Your product lifecycle is already modeled by Kubernetes CRDs. - You need day-2 operations such as scale, stop, start, backup, restore, or delete-backup. - You want Omnistrate to expose those operations through the Customer Portal, APIs, and operations workflows. - You want to keep using standard Kubernetes, Argo Workflow-style YAML, Helm charts, and optionally Terraform. If your application is only a basic container deployment, start with [Build from Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md). If it is a standard Helm chart without custom resources, start with [Build from Helm charts](https://docs.omnistrate.com/getting-started/build-from-helm/index.md). ## Operator Service Plan Flow 1. **Install the operator and CRDs** Determine whether the operator is cluster-scoped or namespace-scoped. Install a cluster-scoped operator and its CRDs once per deployment cell as a [custom amenity](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md). For a namespace-scoped operator, install its CRDs as a deployment-cell amenity and install the controller as a per-instance sibling service with `helmChartConfiguration` and `crds.enabled: false`. Do not add new operator chart entries under `operatorCRDConfiguration.helmChartDependencies`; that pattern is retained only by older examples. 1. **Define customer inputs** Add `apiParameters` for values such as instance type, storage size, replica count, database name, and backup storage settings. 1. **Expose endpoints** Add `endpointConfiguration` so customers can find writer, reader, HTTP, metrics, or other service endpoints after the operator creates the underlying Kubernetes services. 1. **Define lifecycle workflows** Define at least `create`, `modify`, and `delete` so Omnistrate can provision, update, and remove the managed resource through standard lifecycle APIs. Add optional `systemWorkflows` such as `start`, `stop`, `backup`, `restore`, and `deleteBackup` only when the service plan supports those operations. 1. **Optionally define provider operations** Use `customWorkflows` only when your product needs operations beyond the standard platform lifecycle APIs, such as repair, compact, diagnostics, or other provider-defined administrative tasks. 1. **Build and release the plan** Build the service plan with `omnistrate-ctl`, then create a test instance from the Customer Portal or CLI. Note Lifecycle behavior belongs in `systemWorkflows`. Use workflow task `successCondition` and `failureCondition` for readiness and failure handling, and define workflow-level outputs only when a successful task captures the resource status. Operator-created pods may also need explicit affinity to Omnistrate-managed nodes; the operator does not automatically inherit Omnistrate's placement settings. ## Quick Start from the Template ``` git clone https://github.com/omnistrate-community/operator-spec-template cd operator-spec-template omnistrate-ctl build -f spec.yaml --name "Postgres Operator" --release-as-preferred --spec-type ServicePlanSpec ``` The template includes: - CloudNativePG and Barman Cloud plugin Helm dependencies. - `apiParameters` for PostgreSQL credentials, replica count, storage size, instance type, and S3 backup settings. - Public writer and reader endpoint configuration. - `capabilities.backupConfiguration` for periodic backups, retention, and snapshot-before-delete. - `systemWorkflows.create` and `modify` to apply the CNPG `Cluster`. - `systemWorkflows.start` and `stop` to toggle CNPG hibernation. - `systemWorkflows.addCapacity` and `removeCapacity` to change replica count. - `systemWorkflows.backup`, `restore`, and `deleteBackup` to integrate Omnistrate snapshots with operator backup resources. Note The public template uses `operatorCRDConfiguration.helmChartDependencies` for compatibility with its existing installation flow. For new integrations, use deployment-cell amenities for cluster-scoped operators or the namespace-scoped hybrid pattern described above. Deprecated fields such as `operatorCRDConfiguration.template`, `supplementalFiles`, and `readinessConditions` are intentionally omitted; lifecycle resources and readiness checks are modeled with `systemWorkflows`. ## Workflow Syntax Operator lifecycle workflows use an Argo Workflow-style structure: ``` systemWorkflows: create: workflow: entrypoint: create arguments: parameters: - name: namespace value: "{{ $sys.namespace }}" - name: instanceId value: "{{ $sys.instanceId }}" templates: - name: create dag: tasks: - name: applycluster template: apply-cluster - name: apply-cluster resource: action: apply successCondition: status.conditions.#(type=="Ready").status == True failureCondition: status.phase == failed manifest: | apiVersion: postgresql.cnpg.io/v1 kind: Cluster metadata: name: "{{inputs.parameters.instanceId}}" namespace: "{{inputs.parameters.namespace}}" ``` The workflow body uses familiar Argo concepts: `entrypoint`, `arguments.parameters`, `templates`, DAG tasks, and Kubernetes `resource` templates. Omnistrate renders `$sys.*`, `$var.*`, `$secret.*`, and `$func.*` expressions before execution. For the detailed operator service plan model, see [Build with Kubernetes Operators](https://docs.omnistrate.com/build-guides/operators/index.md). For all supported service spec fields, see the [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). ## Validate Your First Instance After you create an instance, check: - The operator Helm charts are installed. - The tenant namespace exists. - The service plan defines the required `systemWorkflows.create`, `systemWorkflows.modify`, and `systemWorkflows.delete` lifecycle hooks. - The custom resources created by `systemWorkflows.create` exist in the tenant namespace. - The live custom resource reaches the operator-specific ready condition. - The operator writes the status fields used by `successCondition`, `failureCondition`, and output parameters. - The Customer Portal shows the expected endpoints and supported operations. - Create, modify, stop, start, and delete have been exercised where the operator supports them, and deletion leaves no custom resources or secrets orphaned in the tenant namespace. - Manual backup and restore work before enabling automated backup schedules. If readiness stalls or a workflow fails, open the workflow details in Operations Center and use [Deployment Cell Access](https://docs.omnistrate.com/operate-guides/deployment-cell-access/index.md) to inspect the Kubernetes resources directly. ## More Real-World Examples - [CloudNativePG Operator PostgreSQL deployment](https://github.com/omnistrate-community/postgres-paas/tree/main/operator) - [Redis Operator](https://github.com/omnistrate-community/examples/tree/main/docs/redis-operator) - [Redis Operator with custom amenities](https://github.com/omnistrate-community/examples/tree/main/docs/redis-operator-customameneties) ## Dive Deeper Now that you have a working Operator-based SaaS Product, explore the build guides to customize and extend your deployment: - [Kubernetes Operators deployment strategy](https://docs.omnistrate.com/build-guides/operators-overview/index.md) — Browse the Operator-specific strategy and guide set - [Operator Troubleshooting](https://docs.omnistrate.com/build-guides/operators-troubleshooting/index.md) — Diagnose failed operator reconciliation and workflows - [Plan Specification Reference](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) — Complete reference for the Plan spec format, including the Operator CRD configuration schema - [Helm Chart Overview](https://docs.omnistrate.com/build-guides/helm-charts-overview/index.md) — Customize Helm charts; use deployment-cell amenities for shared operators and the operator guide for per-instance controller charts - [Helm Chart Customization](https://docs.omnistrate.com/build-guides/helm-charts-customize/index.md) — Customize chart values and affinity rules for your Operator's Helm dependencies - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Expose customer-configurable parameters through your SaaS APIs - [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) — Use dynamic system-generated values in your CRD templates - [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) — Manage shared Operator installations at the cluster level - [Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md) — Set up dependencies between resources for multi-component applications # Building your SaaS Product from Source Repository ## Getting started Omnistrate allows you to build your SaaS Product directly from your source code repository. This approach is ideal when you have application source code in Git repositories with a Dockerfile and want Omnistrate to handle the build and deployment process. ## Overview When you build from a source repository, Omnistrate: - Builds Docker containers from your repository using your existing Dockerfile - Pushes the built image to a container registry - Creates and deploys your service on the Omnistrate platform - Integrates with your CI/CD workflows for automated deployments - Supports any application that can be containerized with Docker This approach is perfect for: - **Applications with Dockerfiles** that are ready to be containerized - **Teams using AI coding assistants** like GitHub Copilot or Cursor to generate Dockerfiles - **Existing containerized applications** that you want to deploy as SaaS - **CI/CD integration** where builds trigger automatically from code changes ## Prerequisites Before building from your repository, ensure you have: 1. **Omnistrate Account**: Sign up at [omnistrate.cloud](https://omnistrate.cloud) 1. **Omnistrate CLI**: Install from [ctl.omnistrate.cloud/install](https://ctl.omnistrate.cloud/install/) 1. **Source Repository**: Your application code in a Git repository (GitHub, GitLab, Bitbucket, etc.) 1. **Dockerfile**: A valid Dockerfile in your repository root (can be generated using AI tools like GitHub Copilot or Cursor) 1. **Docker Daemon**: Docker running on your local machine for building the image 1. **Repository Access**: Ensure you have access to push to container registries ## Supported Applications Since the build-from-repo command works with any valid Dockerfile, it supports any application that can be containerized: ### Web Applications - **Node.js** (Express, Next.js, Nuxt.js, etc.) - **Python** (Django, Flask, FastAPI, etc.) - **Java** (Spring Boot, Maven, Gradle projects) - **Go** applications - **PHP** (Laravel, Symfony, etc.) - **Ruby** (Rails, Sinatra, etc.) - **.NET** applications - **Static sites** (React, Vue, Angular, etc.) ### Databases and Data Services - **PostgreSQL** with custom extensions - **MySQL** with custom configurations - **Redis** with custom modules - **MongoDB** with custom setups - **ClickHouse**, **TimescaleDB**, and other specialized databases ### Microservices and APIs - **REST APIs** in any supported language - **GraphQL** services - **gRPC** services - **Message queues** and event streaming - **Background workers** and batch processing jobs ### AI/ML Applications - **Python ML models** (TensorFlow, PyTorch, scikit-learn) - **Jupyter notebooks** as services - **Model inference APIs** - **Data processing pipelines** ## Getting Started ### Step 1: Prepare Your Repository Ensure your repository contains: 1. **Application Source Code**: Your main application files 1. **Dockerfile**: A valid Dockerfile that can build your application (required) 1. **Dependency Files**: Package files like `package.json`, `requirements.txt`, `pom.xml`, `go.mod`, etc. 1. **Configuration Files**: Any necessary config files for your application #### Creating a Dockerfile with AI Assistance If you don't have a Dockerfile, you can easily create one using AI coding assistants: **Using GitHub Copilot:** 1. Create a new file named `Dockerfile` in your repository root 1. Add a comment describing your application: `# Dockerfile for Node.js Express application` 1. Let Copilot suggest the Dockerfile content **Using Cursor:** 1. Open your project in Cursor 1. Create a `Dockerfile` and use Cursor's AI to generate the content 1. Describe your application stack and let the AI create the appropriate Dockerfile **Example AI prompt:** ``` Create a Dockerfile for a Node.js Express application that: - Uses Node.js 18 Alpine base image - Installs dependencies from package.json - Exposes port 3000 - Runs the application with npm start ``` ### Step 2: Build Your Service Use the Omnistrate CLI to build your service directly from the repository. Run this command from the root of your repository: ``` omnistrate-ctl build-from-repo --product-name "My App" ``` You can also specify additional options: ``` # Build with environment variables (when no compose spec exists) omnistrate-ctl build-from-repo --product-name "My App" --env-var POSTGRES_PASSWORD=default ``` The CLI will: 1. **Build the Docker image** from your Dockerfile 1. **Push the image** to Github Container registry 1. **Generate a compose specification** (if one doesn't exist) 1. **Create your service** on the Omnistrate platform 1. **Set up the infrastructure** according to your specification 1. **Initialize the SaaS portal** for customer access ### Step 4: Monitor the Build Process During the build process, you'll see output similar to: ``` ✓ Docker image built successfully ✓ Container image pushed to registry ✓ Compose specification generated ✓ Service created successfully ✓ Environment promoted to production ✓ SaaS portal initialized Check your service at: https://omnistrate.cloud/product-tier?serviceId=s-abc123&environmentId=se-def456 Access your Customer Portal at: https://saasportal.instance-xyz.omnistrate.cloud/service-plans?serviceId=s-abc123 ``` You can use command line or the customer portal to create an instance of the service. ## Next Steps After successfully building from your repository: 1. **Test Your Service**: Use the provided Customer Portal URL to test deployments 1. **Configure Environments**: Set up development, staging, and production environments 1. **Set Up Monitoring**: Enable logging and metrics 1. **Implement CI/CD**: Automate deployments with your preferred CI/CD platform If you need help setting up your SaaS Product, reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com). ## Dive Deeper Explore the build guides to add production-grade capabilities to your container-based SaaS Product: - [Build Guide Overview](https://docs.omnistrate.com/build-guides/overview/index.md) — Understand the core concepts for building on Omnistrate - [Compose Specification Reference](https://docs.omnistrate.com/build-guides/compose-spec/index.md) — Use Omnistrate Compose extensions to configure your container deployment - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Let your customers configure deployments with custom parameters - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) — Configure hosted, BYOC, or air-gapped deployment models - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) — Choose between shared, dedicated, or hybrid tenancy - [Action Hooks](https://docs.omnistrate.com/build-guides/actionhooks/index.md) — Run custom logic during provisioning, scaling, or upgrades - [App Runtime Guide](https://docs.omnistrate.com/runtime-guides/overview/index.md) — Configure autoscaling, backups, sidecars, and more ## Example Reference For a complete working example, see the [Ray cluster repository](https://github.com/omnistrate-community/ray-cluster) which demonstrates building a distributed computing platform with multiple services and [Private AI Chatbot](https://github.com/omnistrate-community/private-ai-chatbot) which demonstrates multi-tenant ChatBot for managing AI-powered chat interactions. # Building your SaaS Product using Terraform ## Getting started The Omnistrate Terraform Integration streamlines the deployment and management of cloud infrastructure by combining the power of Terraform with the flexibility of Omnistrate. This integration allows you to inject dynamic metadata directly into your Terraform templates, automate complex multi-cloud deployments, and modularize infrastructure components for efficient management. By integrating with other deployment types such as Helm charts, Operators or Kustomize, Omnistrate enhances orchestration capabilities, enabling users to create scalable and adaptable Plans. With Omnistrate, you can simplify your infrastructure-as-code workflows and ensure smooth, automated provisioning across various cloud environments and tenants. Note Support for Terraform Integration is available for AWS, GCP, Azure, OCI, and Nebius. You can define your Plan to use the corresponding stack depending on the cloud provider that is being deployed, making your service multi-cloud. Warning Omnistrate manages the Terraform state backend using Kubernetes secrets. Any custom `backend` configuration in your Terraform files (such as S3, GCS, or Azure Blob Storage) will be removed during deployment. Do not include a custom backend block — Omnistrate handles state management automatically. ## Integrating Terraform on your Plan Terraform stacks require a Plan specification file that defines your overall service topology on Omnistrate. Unlike Docker Compose deployments where the service specification can be embedded within the compose file, Terraform deployments must use a separate Plan specification file. A complete description of the Plan specification can be found on [Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) Here is an example of a using Terraform to configure a SaaS Product: ``` provider "aws" { region = "{{ $sys.deploymentCell.region }}" } # Create a Security Group for RDS and ElastiCache resource "aws_security_group" "rds_elasticache_sg" { name = "e2e-rds-elasticache-security-group-{{ $sys.id }}" description = "Security group for RDS and ElastiCache instances" vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" ingress { from_port = 3306 to_port = 3306 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] # Adjust for appropriate security } ingress { from_port = 11211 # Default Memcached port to_port = 11211 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] # Adjust accordingly } egress { from_port = 0 to_port = 0 protocol = "-1" cidr_blocks = ["0.0.0.0/0"] } } # Create a DB Subnet Group for RDS resource "aws_db_subnet_group" "rds_subnet_group" { name = "e2e-rds-subnet-group-{{ $sys.id }}" description = "My RDS subnet group" subnet_ids = [ "{{ $sys.deploymentCell.publicSubnetIDs[0].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[1].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[2].id }}" ] } # Create a Subnet Group for ElastiCache resource "aws_elasticache_subnet_group" "elasticache_subnet_group" { name = "e2e-elasticache-subnet-group-{{ $sys.id }}" description = "My ElastiCache subnet group" subnet_ids = [ "{{ $sys.deploymentCell.publicSubnetIDs[0].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[1].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[2].id }}" ] } # Create RDS instances resource "aws_db_instance" "example1" { identifier = "e2e-instance-1-{{ $sys.id }}" engine = "mysql" instance_class = "db.t3.micro" allocated_storage = 20 db_subnet_group_name = aws_db_subnet_group.rds_subnet_group.name vpc_security_group_ids = [aws_security_group.rds_elasticache_sg.id] username = "admin" password = "yourpassword" parameter_group_name = "default.mysql8.0" engine_version = "8.0.37" skip_final_snapshot = true depends_on = [ aws_security_group.rds_elasticache_sg, aws_db_subnet_group.rds_subnet_group ] } resource "aws_db_instance" "example2" { identifier = "e2e-instance-2-{{ $sys.id }}" engine = "mysql" instance_class = "db.t3.micro" allocated_storage = 20 db_subnet_group_name = aws_db_subnet_group.rds_subnet_group.name vpc_security_group_ids = [aws_security_group.rds_elasticache_sg.id] username = "admin" password = "yourpassword" parameter_group_name = "default.mysql8.0" engine_version = "8.0.37" skip_final_snapshot = true depends_on = [aws_db_instance.example1] } # Create ElastiCache Cluster for Memcached resource "aws_elasticache_cluster" "example_memcached" { cluster_id = "e2e-memcached-{{ $sys.id }}" engine = "memcached" node_type = "cache.t3.micro" num_cache_nodes = 2 # Adjust as needed subnet_group_name = aws_elasticache_subnet_group.elasticache_subnet_group.name security_group_ids = [aws_security_group.rds_elasticache_sg.id] depends_on = [ aws_db_instance.example2 # Ensure ElastiCache is created after all RDS instances ] } ``` Warning Please ensure that the main file for your terraform stack contains the definition on the provider. Terraform templates can be referenced from a GitHub repository, either private or public. One repository can contain multiple Terraform stacks, and can be referenced using a specific git reference (tag or branch) and a folder path within that repository. Omnistrate allows you to configure a reference to a Git repository and path where the Terraform stack is stored. When selecting a path for your repository, Omnistrate expects the entire Terraform definition to be under that folder structure. You can also use a local Terraform artifact with `omnistrate-ctl build`. Use `artifactRelativePath` in the Plan spec and run CTL from the workspace root that contains that relative path. Example workspace: ``` my-service/ spec.yaml terraform/ aws/ main.tf variables.tf outputs.tf ``` Example local-artifact spec: ``` name: Terraform Local Artifact deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: "arn:aws:iam:::role/omnistrate-bootstrap-role" services: - name: terraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: / artifactRelativePath: terraform/aws ``` Build the Plan with CTL: ``` cd my-service omnistrate-ctl build \ --spec-type ServicePlanSpec \ --file spec.yaml \ --product-name "Terraform Local Artifact" \ --release-as-preferred ``` There are two paths to configure: - `artifactRelativePath` tells CTL which local file or directory to upload. - `terraformPath` tells Omnistrate where to run OpenTofu or Terraform inside that uploaded content. For example, if your Terraform module is in `terraform/aws`, you can upload that module directory directly: ``` artifactRelativePath: terraform/aws terraformPath: / ``` Or you can upload its parent directory and point `terraformPath` at the module subdirectory: ``` artifactRelativePath: terraform terraformPath: /aws ``` Use one Terraform source per cloud-provider entry: set either `gitConfiguration` or `artifactRelativePath`, not both. Prefer relative artifact paths such as `terraform/aws`; paths that escape the workspace root, such as `../terraform`, are rejected. Within the Terraform templates, the Omnistrate platform provides system parameters that can be used to reference to information about the current cluster, for instance: - `$sys.deploymentCell.publicSubnetIDs[i].id` - `$sys.deploymentCell.privateSubnetIDs[i].id` - `$sys.deploymentCell.region` - `$sys.deploymentCell.cloudProviderNetworkID` - `$sys.deploymentCell.cidrRange` You can inject these values into your Terraform templates, and these values will be dynamically rendered during deployment. In addition, the Omnistrate platform will automatically output all values defined in the Terraform templates' output block. You can use these outputs to inject into other deployment components, as shown in the example below: Info You can use system parameters to customize Terraform templates. A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ## Registering a Service using a Terraform stack Terraform stacks are managed through a specification file that defines your overall service topology on Omnistrate. A complete description of the Plan specification can be found on [Getting started / Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). Consider using Git tags to version the terraform stack and ensure consistency across Plan versions. Here is an example of a SaaS Product that deploys Redis Clusters using a Helm chart. Note This example uses the legacy `operatorCRDConfiguration.template` and `readinessConditions` fields to show Terraform outputs flowing into an operator resource. For new operator-backed services, prefer `systemWorkflows` for lifecycle resources, readiness, failure conditions, and outputs. See [Build from Kubernetes Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md). **Example:** ``` name: Terraform Resources deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: terraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /artifacts/terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/omnistrate-community/examples.git - name: CNPG dependsOn: - terraform operatorCRDConfiguration: template: | apiVersion: postgresql.cnpg.io/v1 kind: Cluster metadata: name: cluster-example namespace: default # Should ignore this namespace spec: instances: 3 storage: size: 1Gi readinessConditions: "$var._crd.status.phase": "Cluster in healthy state" "$var._crd.status.readyInstances": 3 '$var._crd.status.conditions[?(@.type=="Ready")].status': "True" outputParameters: "Postgres Container Image": "$var._crd.status.image" "Status": "$var._crd.status.phase" "Topology": "$var._crd.status.topology" helmChartDependencies: - chartName: cloudnative-pg chartVersion: 0.22.1 chartRepoName: cnpg chartRepoURL: https://cloudnative-pg.github.io/charts ``` You can register this spec using our CLI: ``` omnistrate-ctl build -f spec.yaml --name 'Multiple Resources' --release-as-preferred --spec-type ServicePlanSpec # Example output shown below ✓ Successfully built service Check the Plan result at: https://omnistrate.cloud/product-tier?serviceId=s-xxx&environmentId=se-xxx Access your SaaS Product at: https://saasportal.instance-xxx.hc-xxx.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-xxxx&environmentId=se-xxx ``` ## More Real-World Examples - [AWS RDS Aurora PostgreSQL deployment](https://github.com/omnistrate-community/postgres-paas/tree/main/paas) - [Terraform infrastructure services](https://github.com/omnistrate-community/examples/tree/main/docs/terraform-helm#terraform-infrastructure-services) ## Using Terraform outputs Omnistrate platform will automatically output all values defined in the Terraform templates output block. You can use these outputs to inject into other deployment components. For example, a resource can define output parameters on the stack ``` output "rds_endpoints_1" { value = aws_db_instance.example1.endpoint } output "rds_endpoints_2" { value = { endpoint: aws_db_instance.example2.endpoint } sensitive = true } output "elasticache_endpoint" { value = aws_elasticache_cluster.example_memcached.cache_nodes[0].address } ``` and in the dependent resource we can reference to the terraform output properties: ``` {{ $terraformChild.out.rds_endpoints_1 }} {{ $terraformChild.out.rds_endpoints_2.endpoint }} {{ $terraformChild.out.elasticache_endpoint }} ``` If you want selected Terraform outputs to appear as exported fields on the Terraform resource in Omnistrate, declare them explicitly under `requiredOutputs`: ``` services: - name: terraformChild internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git requiredOutputs: - key: rds_endpoints_1 exported: true - key: elasticache_endpoint exported: true ``` For a deeper explanation of output mapping versus exported Terraform outputs, see [Input Parameters and Output Mapping](https://docs.omnistrate.com/build-guides/terraform-params-outputs/index.md). ## Additional permissions for Terraform Using a terraform stack normally requires to define permissions to create, update and delete the entities involved. You can configure a custom policy / role that will be used when applying each Terraform stack defined as a resource in Omnistrate. ### Custom Terraform permissions for BYOA Deployment Model To configure permissions required to provision Terraform resources within BYOA Deployment Model, you can configure a Custom Terraform Policy. This feature is configured on the Plan level and Omnistrate will adjust onboarding scripts for your customers to include appropriate resources (depending on cloud provider). - For AWS, configuration is provider as a policy document. Omnistrate will create a dedicated AWS IAM role per-model when onboarding your customers. With each update, a new version of the cloud formation stack will be created. - For GCP, configuration is provided as a list of [GCP IAM roles](https://docs.cloud.google.com/iam/docs/roles-permissions/). Omnistrate will create a dedicated GCP IAM service account per-role with specified roles bound. GCP CLI commands will always reflect the latest configuration. - For Azure, configuration is provided as a list of [RBAC roles](https://learn.microsoft.com/en-us/azure/role-based-access-control/built-in-roles). Omnistrate will create a single service principal with permissions merged across all service models. RBAC permissions are scoped to the subscription level. [Entra roles](https://learn.microsoft.com/en-us/entra/identity/role-based-access-control/permissions-reference) can be bound by specifying the type as "Entra" and are bound to the root tenant level. Just like for the GCP, an Azure CLI onboarding script will always reflect the latest configuration. - For OCI, configuration is provided as a list of OCI IAM permission fragments under `permissions.oci`, such as `manage queues` and `manage ons-topics`. During onboarding, Omnistrate creates a Terraform user and renders these fragments into policy statements scoped to the account compartment created for the account config: `Allow any-user to in compartment id where request.principal.id=''`. During provisioning, this custom identity (AWS IAM role / GCP IAM service account) is used to modify Terraform resources within the Plan. To configure custom terraform permissions using the service spec file, add the "CUSTOM_TERRAFORM_POLICY" feature such as: ``` features: CUSTOM_TERRAFORM_POLICY: policies: aws: | { "Statement": [ { "Action": [ "sqs:*" ], "Effect": "Allow", "Resource": "*" } ] } roles: gcp: - name: roles/pubsub.admin - name: roles/cloudsql.admin azure: - name: "Network Contributor" - name: "Storage Account Contributor" - name: "Virtual Machine Contributor" - name: "Global Reader" type: "Entra" permissions: oci: - "manage queues" - "manage ons-topics" ``` In the example above, Terraform will be allowed to manage OCI Queues and OCI Notifications (ONS) topics. To achieve the same effect using Omnistrate CTL, you can use the following command: ``` omctl service-plan enable-feature 'product name' 'Plan name' --feature CUSTOM_TERRAFORM_POLICY --feature-configuration-file ../path/to/config-file ``` With config file having the following content (AWS policy itself needs to be provided as JSON string): ``` { "policies": { "aws": "{\"Statement\":[{\"Action\": [\"sqs:*\"],\"Effect\": \"Allow\",\"Resource\": \"*\"}]}" }, "roles": { "gcp": [ {"name": "roles/pubsub.admin"}, {"name": "roles/cloudsql.admin"} ], "azure": [ {"name": "Network Contributor"}, {"name": "Storage Account Contributor"}, {"name": "Virtual Machine Contributor"}, {"name": "Global Reader", "type": "Entra"} ] }, "permissions": { "oci": [ "manage queues", "manage ons-topics" ] } } ``` To disable this feature, use the following command: ``` omctl service-plan disable-feature 'product name' 'Plan name' --feature CUSTOM_TERRAFORM_POLICY ``` Warning Each time the custom terraform policy is updated, a new version of the onboarding script is generated. Your customers might need to re-run their account onboarding, otherwise they might not be able to provision your service. This requires updating their account onboarding CloudFormation stack (for AWS), or getting latest GCP CLI commands and running them (for GCP). ### Custom Terraform execution identity for Hosted SaaS For Hosted SaaS deployments, Omnistrate manages Terraform execution identities for each cloud provider. For AWS, GCP, Azure, and OCI, the identity is auto-created — you only need to assign the permissions your Terraform stack requires. You do not need to specify the identity in the Plan spec. - **AWS** (auto-created): Omnistrate creates an IAM role named `omnistrate-terraform-execution-role` in your AWS account. Assign the permissions your Terraform stack needs to this role. - **GCP** (auto-created): Omnistrate creates a service account named `omnistrate-tf-` in your GCP project. Assign the permissions your Terraform stack requires to this service account. - **Azure** (auto-created): Omnistrate creates a service principal named `terraform--`. Assign the permissions your Terraform stack requires to this service principal. - **OCI** (auto-created): Omnistrate creates a user named `-terraform-user`. Assign the permissions your Terraform stack requires to this user. ``` name: Multiple Resources deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: terraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/omnistrate-community/examples.git gcp: terraformPath: /terraform gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/omnistrate-community/examples.git ``` ### Nebius Terraform authentication Nebius uses explicit service-account authentication. Configure the service-account credentials on the `nebius` Terraform entry: ``` services: - name: terraform internal: true terraformConfigurations: configurationPerCloudProvider: nebius: terraformPath: /terraform/nebius serviceAccountID: "serviceaccount-e00vqdp9fskhmmaan8" publicKeyID: "publickey-e00h9scsyy9mbefrjf" privateKeyPEM: "$secret.nebiusTerraformPrivateKey" gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Notes: - `serviceAccountID`, `publicKeyID`, and `privateKeyPEM` are all required for `configurationPerCloudProvider.nebius`. - `privateKeyPEM` can be inline, but prefer `$secret.` so the PEM stays out of source control. - Omnistrate resolves `$secret` references before it hands the stack to Terraform. - Keep the Terraform source provider block minimal. Omnistrate injects the Nebius credential wiring at runtime; for example: ``` terraform { required_providers { nebius = { source = "terraform-provider.storage.eu-north1.nebius.cloud/nebius/nebius" } } } provider "nebius" { domain = "api.eu.nebius.cloud:443" } ``` - When Nebius auth is specified, Omnistrate uses it for that Terraform resource instead of the default host-cluster Terraform identity. ### Control-Plane-Targeted Terraform Some Terraform resources manage provider-side assets that should run in the Omnistrate control plane account instead of the data plane deployment cell. For those resources, set `deploymentTarget.account: ControlPlane` on the Terraform resource. Example: ``` services: - name: controlPlaneInfra internal: true deploymentTarget: account: ControlPlane terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/control_plane gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Note For control-plane-targeted Terraform, ensure that the appropriate execution identity exists in your cloud account. For AWS, GCP, Azure, and OCI, Omnistrate auto-creates the identity. Assign the permissions your Terraform stack requires to the identity. Notes: - For control-plane-targeted Terraform, define only the cloud provider entry that actually executes the stack. A control-plane-targeted AWS stack does not need placeholder `gcp`, `azure`, or `nebius` entries. - There is no shared shorthand outside `configurationPerCloudProvider`; each provider entry is defined explicitly. - This is useful for shared control-plane resources such as DNS, registry, or other provider-side integrations that are not created inside each deployment cell. #### AWS Terraform role Omnistrate pre-creates the `omnistrate-terraform-execution-role` IAM role for Terraform executions during account setup. You do not need to create this role for AWS. ##### AWS requirements for Terraform role The role uses the exact name `omnistrate-terraform-execution-role`. Omnistrate creates the role with the required Trusted Entity configuration: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam:::root" }, "Action": "sts:AssumeRole", "Condition": { "StringLike": { "aws:PrincipalArn": "arn:aws:iam:::role/dataplane-agent-iam-role-*" } } } ] } ``` Info Omnistrate creates the IAM role with the exact name `omnistrate-terraform-execution-role`. #### Pre-created principals Omnistrate will by default pre-create empty IAM principals for Terraform resources during account setup. - For AWS, IAM role `omnistrate-terraform-execution-role` will be created (this role starts with `omnistrate-` and has a trusted entity configuration) - For GCP, IAM service account `omnistrate-tf-` will be created (this service account will have IAM binding to Omnistrate dataplane) - For Azure, service principal `terraform--` will be created - For OCI, user `-terraform-user` will be created These principals are pre-created with necessary configuration (as described above). You can configure them with the necessary permissions for your Terraform stack. This pre-created-principal flow does not apply to Nebius Terraform resources; use the Nebius service-account fields shown above. ## Terraform Variables Override with Inline File Omnistrate supports overriding Terraform variables using inline file content through the `variablesValuesFileOverride` property. This feature allows you to define Terraform variable values directly within your Plan specification without requiring a separate `.tfvars` file in your repository. The `variablesValuesFileOverride` property accepts a multi-line string containing Terraform variable definitions in the standard `.tfvars` format. You can use Omnistrate system parameters within these variable definitions to inject dynamic values during deployment. ### Configuration Example Here's an example of how to use the `variablesValuesFileOverride` feature: ``` name: Multiple Resources deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' services: - name: terraformChild internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform variablesValuesFileOverride: | vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" instance_type = "t3.medium" environment = "production" subnet_ids = [ "{{ $sys.deploymentCell.publicSubnetIDs[0].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[1].id }}" ] gitConfiguration: reference: refs/tags/12.0 repositoryUrl: https://github.com/omnistrate-community/sample/TestKustomizeTemplate.git ``` ### Using System Parameters You can inject Omnistrate system parameters into your variable values: ``` variablesValuesFileOverride: | # Network configuration from deployment cell vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" availability_zones = [ "{{ $sys.deploymentCell.region }}a", "{{ $sys.deploymentCell.region }}b" ] # Dynamic naming with system ID resource_prefix = "omnistrate-{{ $sys.id }}" # Subnet configuration public_subnets = [ "{{ $sys.deploymentCell.publicSubnetIDs[0].id }}", "{{ $sys.deploymentCell.publicSubnetIDs[1].id }}" ] private_subnets = [ "{{ $sys.deploymentCell.privateSubnetIDs[0].id }}", "{{ $sys.deploymentCell.privateSubnetIDs[1].id }}" ] ``` Tip The `variablesValuesFileOverride` feature is particularly useful when you want to customize Terraform behavior on Omnistrate platform without modifying the underlying Terraform code. ## Custom OpenTofu CLI Configuration Omnistrate runs your Terraform stack with OpenTofu. Use the `cliConfigFileOverride` property to supply a custom [OpenTofu CLI configuration file](https://opentofu.org/docs/cli/config/config-file/) for a Terraform Resource — without baking it into your repository or a custom runner image. The property accepts a multi-line string containing the CLI configuration file content. Omnistrate writes it into the Terraform workspace and points `TF_CLI_CONFIG_FILE` at it for every operation on the stack (`init`, `plan`, `apply`, `destroy`) for the cloud provider entry where you specify it. It works for both Git-based and artifact-based Terraform sources. Common uses include provider installation mirrors and credentials for private module registries. ### Example: Alternate Provider Mirrors for Specific Registries The following configuration installs `hashicorp/*` providers from an internal network mirror while all other providers continue to install directly from their origin registries — useful when the public registry is unreachable from your deployment environment, or when provider binaries must come from a vetted internal source: ``` services: - name: terraformResource internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform cliConfigFileOverride: | provider_installation { network_mirror { url = "https://terraform-mirror.internal.example.com/providers/" include = ["registry.opentofu.org/hashicorp/*"] } direct { exclude = ["registry.opentofu.org/hashicorp/*"] } } gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` With this configuration, `tofu init` resolves any provider matching `registry.opentofu.org/hashicorp/*` through the mirror at `terraform-mirror.internal.example.com`, and falls back to direct installation for everything else. See the [OpenTofu CLI configuration documentation](https://opentofu.org/docs/cli/config/config-file/) for the full set of supported blocks, including `credentials` for private registries and `oci_default_credentials` for OCI-based mirrors. Tip Like `variablesValuesFileOverride`, the CLI configuration content supports Omnistrate system parameters, so mirror URLs or registry tokens can be injected dynamically at deployment time. Note CLI configuration files can contain credentials. Omnistrate stores the content on the Terraform Resource and redacts `cliConfigFileOverride` in API responses for users with read-only roles. ## Dive Deeper Now that you have a working Terraform-based SaaS Product, explore the build guides to extend your deployment: - [Terraform deployment strategy](https://docs.omnistrate.com/build-guides/terraform-overview/index.md) — Browse the Terraform-specific strategy and guide set - [Terraform Troubleshooting](https://docs.omnistrate.com/build-guides/terraform-troubleshooting/index.md) — Diagnose failed Terraform plans and applies - [Terraform Multi-Cloud Configuration](https://docs.omnistrate.com/build-guides/terraform-multi-cloud/index.md) — Configure Terraform stacks across AWS, GCP, Azure, OCI, and Nebius - [Terraform Input Parameters and Output Mapping](https://docs.omnistrate.com/build-guides/terraform-params-outputs/index.md) — Map API parameters to Terraform variables and expose outputs - [Helm and Terraform](https://docs.omnistrate.com/build-guides/helm-charts-terraform/index.md) — Combine Terraform-managed infrastructure with Helm-deployed applications - [Plan Specification Reference](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) — Complete reference for the Plan spec format, including the Terraform configuration schema - [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md) — Expose customer-configurable parameters through your SaaS APIs - [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) — Inject dynamic platform values into your Terraform variables - [Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md) — Set up dependencies between Terraform and other resources (Helm, Operators) # Quick start with Command Line ## What is Omnistrate CTL? Omnistrate CTL is a command line tool that helps you build SaaS Products, manage Plans, and operate instance deployments efficiently. Install the CLI before you continue by following [Installing Omnistrate CTL](https://docs.omnistrate.com/getting-started/installing-ctl/index.md). ## Quick start video **Note:** Some features may have been updated since the video was created. Refer to the [omnistrate-ctl manual](https://ctl.omnistrate.cloud/omnistrate-ctl/) for the most up-to-date information. If you are onboarding hosted cloud accounts or customer BYOA accounts from the CLI, start with the [Cloud Account Guides](https://docs.omnistrate.com/getting-started/onboarding/overview/index.md). ### Build Service from Compose Spec Before you can build a SaaS Product, you need a Compose specification that defines the product. The Compose specification file should include the Plan configuration in the [x-omnistrate-service-plan](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-service-plan) section. For more information on creating a Compose specification, refer to the [Compose spec reference](https://docs.omnistrate.com/build-guides/compose-spec/index.md). To build a service from a compose spec file, use the `build` command with the following information: - `--file`: The path to the compose spec file (optional will default to omnistrate-compose.yaml, docker-omnistrate-compose.yaml, or spec.yaml) - `--product-name`: The name of the service to create. - `--description`(Optional): A description of the service. - `--service-logo-url`(Optional): The URL of the service logo. - `--release`(Optional): Flag to release the service. This will make the Plan available to create instance deployments. - `--release-as-preferred`(Optional): Flag to release the service as preferred. This will make the service the default version used to create instance deployments. - `--environment`(Optional): The name of the environment where the service will be built. Default is the dev environment. - `--environment-type`(Optional): The type of environment where the service will be built. Use together with the `--environment` flag. Default is the dev environment. - `--interactive`(Optional): Flag to enable interactive mode for building the service. This will prompt you to access your Customer Portal and promoting the Plan to production. ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Your product name" --interactive ``` ### Automation and build behavior notes For CI or other non-interactive workflows, you can log in by piping the password over stdin: ``` cat ~/omnistrate_pass.txt | omnistrate-ctl login --email you@example.com --password-stdin ``` Keep the following build behaviors in mind: - `--environment` must match the configured environment name exactly. If omitted, `build` targets the `Dev` environment. - Use `--environment-type` together with `--environment` so the service plan resolves to the intended environment identity. - When you release a build, include `--release-description` so the resulting Plan version is easier to audit later. - If the rendered spec is unchanged, Omnistrate may reuse the existing Plan version instead of creating a new one. Use `--force-create-service-plan-version` only when you explicitly need a new released version for identical inputs. - If you changed Terraform or Helm inputs and need Omnistrate to re-render artifacts, publish a new Plan version instead of restarting an old workflow. See [Workflows](https://docs.omnistrate.com/operate-guides/workflows/#restarting-a-workflow-vs-publishing-a-new-plan-version) and [Building your SaaS Product using Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md). You can always rerun the `build` command with the same product name to update the service with a new version of the compose spec file. Refer to the [Example](#launch-saas-product-example) for a step-by-step guide on creating a SaaS Product and updating it as you evolve your service. ## Launch SaaS Product Example In this example, we will create a SaaS Product for a Postgres service with a free tier plan, a premium plan, and walk you through Plan updates, releases, and launch to production. You can create instance deployments in your SaaS Product and try it out. After completing this example, you will have a better understanding of how to use the CTL to build and manage your SaaS Products. ### Step 1: Create a SaaS Product with a Free Tier Plan To start, create a free tier plan for the Postgres service. Save `postgres-free-v1.yaml` with the following contents: ``` version: "3.9" x-omnistrate-service-plan: name: 'Postgres Free' tenancyType: 'OMNISTRATE_MULTI_TENANCY' services: postgres: image: postgres ports: - '5432:5432' environment: - SECURITY_CONTEXT_USER_ID=999 - SECURITY_CONTEXT_GROUP_ID=999 - POSTGRES_USER=default - POSTGRES_PASSWORD=default - PGDATA=/var/lib/postgresql/data/dbdata volumes: - ./data:/var/lib/postgresql/data deploy: resources: limits: cpus: '0.50' memory: 50M reservations: cpus: '0.25' memory: 20M x-omnistrate-capabilities: autoscaling: minReplicas: 1 maxReplicas: 5 serverlessConfiguration: enableAutoStop: true minimumNodesInPool: 1 targetPort: 5432 ``` In this example, a free tier plan is created for the Postgres service, optimized for cost savings and scalability. The plan leverages multitenancy and serverless configuration, including auto-stop to minimize costs when the service is not in use. Run the following command to build the service: ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Postgres" --release-as-preferred --interactive ``` Follow the prompts to complete the service creation process, initializing the Customer Portal, promoting the Plan to production, and waiting for production Customer Portal to be ready. Once the Customer Portal is ready, you can view it in your browser and start creating instance deployments. ### Step 2: Enhance your SaaS Product by Adding a Premium Plan Next, let's enhance the service by adding a premium plan. Save `omnistrate-compose.yaml` with the following contents: ``` version: "3.9" x-omnistrate-service-plan: name: 'Postgres Premium' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' services: postgres: image: postgres ports: - '5432:5432' environment: - SECURITY_CONTEXT_USER_ID=999 - SECURITY_CONTEXT_GROUP_ID=999 - POSTGRES_USER=default - POSTGRES_PASSWORD=default - PGDATA=/var/lib/postgresql/data/dbdata volumes: - ./data:/var/lib/postgresql/data x-omnistrate-capabilities: autoscaling: minReplicas: 1 maxReplicas: 5 ``` This plan offers dedicated tenancy with enhanced performance and resource allocation for users who require more robust features and higher limits. Run the following command to build the service and promote it to production following the same steps as before: ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Postgres" --release-as-preferred --interactive ``` ### Step 3: Update your SaaS Product as You Evolve your Service Based on the previous example, let's update the Postgres service's free tier plan to include more features and capabilities. Save `omnistrate-compose.yaml` with the following contents: ``` version: "3.9" x-omnistrate-service-plan: name: 'Postgres Free' tenancyType: 'OMNISTRATE_MULTI_TENANCY' x-customer-integrations: logs: metrics: services: postgres: image: postgres ports: - '5432:5432' environment: - SECURITY_CONTEXT_USER_ID=999 - SECURITY_CONTEXT_GROUP_ID=999 - POSTGRES_USER=default - POSTGRES_PASSWORD=default - PGDATA=/var/lib/postgresql/data/dbdata volumes: - ./data:/var/lib/postgresql/data deploy: resources: limits: cpus: '0.50' memory: 50M reservations: cpus: '0.25' memory: 20M x-omnistrate-capabilities: autoscaling: minReplicas: 1 maxReplicas: 5 serverlessConfiguration: enableAutoStop: true minimumNodesInPool: 1 targetPort: 5432 ``` In this example, the free tier plan for the Postgres service is updated to include logs and metrics integration and enable autoscaling. These enhancements provide better monitoring and scalability for the service. Run the following command to update the service and follow the prompts to promote the new version to production: ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Postgres" --release-as-preferred --interactive ``` ### Step 4: Mounting Configuration and Secret Files (Optional) The CTL allows users to define and manage configuration files and secret files required by their services. These files can be specified in the compose specification and will be automatically mounted to the specified paths in the service containers. #### Configuration Files To specify configuration files in the compose spec file, use the following format: ``` services: service: configs: - source: my_config # The name of the config target: /etc/config/my_config.txt # The target path in the container configs: my_config: # The name of the config file: ./my_config.txt # The path to the config file in your local filesystem ``` This example shows how to define a config named `my_config` and mount it to the path `/etc/config/my_config.txt` within the service container. #### Secret Files Similarly, you can specify secret files as follows: ``` services: service: secrets: - source: server-certificate # The name of the secret target: /etc/ssl/certs/server.cert # The target path in the container secrets: server-certificate: # The name of the secret file: ./server.cert # The path to the secret file in your local filesystem ``` Here is a comprehensive example of a compose spec file for a Postgres service that includes both configuration and secret files: ``` version: "3.9" x-omnistrate-service-plan: name: 'Postgres Premium' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' services: postgres: image: postgres configs: - source: postgres-config target: /etc/postgresql/postgresql.conf secrets: - source: postgres-secret target: /run/secrets/postgres.secret ports: - '5432:5432' environment: - SECURITY_CONTEXT_USER_ID=999 - SECURITY_CONTEXT_GROUP_ID=999 - POSTGRES_USER=default - POSTGRES_PASSWORD=default - PGDATA=/var/lib/postgresql/data/dbdata volumes: - ./data:/var/lib/postgresql/data x-omnistrate-capabilities: autoscaling: minReplicas: 1 maxReplicas: 5 configs: postgres-config: file: ./config/postgres.conf secrets: postgres-secret: file: ./secrets/postgres.secret ``` In this example, the postgres service uses a configuration file for Postgres settings and a secret file for sensitive information. These files are specified in the configs and secrets sections and mounted to the appropriate paths within the container. Name the above file as `omnistrate-compose.yaml` and run the following command to build the service. Make sure the paths to the configuration and secret files are correct and accessible from the location where you run the CTL command. ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "Postgres" --release-as-preferred --interactive ``` # Quick start using Omnistrate Web UI ## Quick start To get started with Omnistrate Web UI, its really simple and you just need to follow quick self-onboarding steps: ### Configure Cloud Accounts To configure your own account, you can follow the steps described in [Account Onboarding](https://docs.omnistrate.com/getting-started/account-onboarding/index.md) ### Build SaaS Product There are several options to start building your SaaS Product: - Click "Hello world project" to build an example Mage AI service - Click "Bring your container image" to build a service from your container image - Click "Bring your Docker Compose" to build a service with your own Compose spec or one of our Certified templates. - Reach out to `support@omnistrate.com` to build a service with Kubernetes Operators, Helm Charts, Terraform, OpenTofu and more or follow our guides on how to onboarding with Omnistrate CTL. Note To learn more on how to build with your Compose spec, please see [here](https://docs.omnistrate.com/getting-started/build-from-compose/index.md). For any questions or help, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) and we will be happy to help. Here is how it looks like: Once you have submitted your service build, Omnistrate will build your service definition and deploy your SaaS Product in a Development environment ### Launch SaaS Product Once your SaaS Product has been created, you can optionally decide to make your SaaS Product live for any one to sign-up and access publicly. Here is how it looks like: If you click "Yes", Omnistrate will deploy your SaaS Product in a Production environment. ## Quick start video ## Hello world example using PostgreSQL One of the easiest ways to get going is to leverage compose templates. You can existing templates or create your own. Here is a sample compose spec to create single writer Postgres SaaS: ``` version: '3.9' services: postgres: image: postgres ports: - '5432:5432' environment: - SECURITY_CONTEXT_USER_ID=999 - SECURITY_CONTEXT_GROUP_ID=999 - POSTGRES_USER=username - POSTGRES_PASSWORD=password - PGDATA=/var/lib/postgresql/data/dbdata volumes: - ./data:/var/lib/postgresql/data ``` In the above case: - We are using only 1 Resource (postgres) with the standard image (postgres) and infra (port: 5432) - Given no account config is specified, the infrastructure for your SaaS is provisioned in the Omnistrate account - There are a few static environment variables that will be injected into the runtime process like POSTGRES_USER In general, the compose config allows you to specify: - Resources and relationship across them - Image(s) - Environment - Infrastructure - Capabilities - Account config information - Enable/Disable Integrations - SaaS experience (input/output params with constraints) For detailed setup instructions, please visit [here](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) Here is the generated Postgres SaaS from the above spec looks like: # Installing Omnistrate Command Line ## What is Omnistrate CTL? Omnistrate CTL is a command line tool designed to streamline the creation, deployment, and management of your Omnistrate SaaS. Use it to build services from images / compose specs, manage Plans, and manage instance deployments efficiently. ## Obtaining CTL Follow the instructions for [omnistrate-ctl installation](https://ctl.omnistrate.cloud/install/) Detailed documentation on the commands available can be found in the [omnistrate-ctl manual](https://ctl.omnistrate.cloud/omnistrate-ctl/) ## Getting Started with CTL To start using CTL, you first need to log in to the Omnistrate platform. You can do this by either providing your email and password or using a single sign-on (SSO) provider. You can choose the login option by running the following command: ``` omnistrate-ctl login ``` For more login methods, refer to the [Login to Omnistrate](https://ctl.omnistrate.cloud/install/#login-to-omnistrate) guide. ## Next Steps To continue, read the [Getting Started with CTL](https://docs.omnistrate.com/getting-started/getting-started-with-ctl/index.md) guide to learn effective usage, or consult the [CTL Command Reference](https://ctl.omnistrate.cloud/omnistrate-ctl/) to explore all available commands and begin building your SaaS with Omnistrate CTL. # Configuring Omnistrate MCP Server Connect your AI tools to Omnistrate using the [Model Context Protocol (MCP)](https://modelcontextprotocol.io), an open standard that lets AI assistants interact with your Omnistrate services and infrastructure. ## What is Omnistrate MCP? Omnistrate MCP is Omnistrate's official MCP server. It is a local MCP server that gives AI tools secure access to your Omnistrate projects and resources through the `omnistrate-ctl` command-line interface. It integrates with popular AI assistants like Claude, enabling you to: - Search and navigate Omnistrate documentation and knowledge base - Manage services, environments, and service plans - Create and manage instances and deployments - Monitor service health and operational events - Manage subscriptions and customer access - Configure deployment cells and cloud accounts - Analyze costs and audit logs ## Connecting to Omnistrate MCP To use Omnistrate MCP, you need to have `omnistrate-ctl` installed and configured on your system. The MCP server communicates with Omnistrate through your local CLI installation. ### Prerequisites - `omnistrate-ctl` CLI tool installed and in your PATH - Omnistrate account with appropriate permissions - Authentication configured for your Omnistrate account ## Supported clients Omnistrate MCP can be integrated with various AI tools and IDEs that support the Model Context Protocol: - [Claude Code](#claude-code) - [Claude.ai and Claude for Desktop](#claudeai-and-claude-for-desktop) - [Cursor](#cursor) - [VS Code with Copilot](#vs-code-with-copilot) - [Cline](#cline) - [Windsurf](#windsurf) - [Gemini Code Assist](#gemini-code-assist) - [Gemini CLI](#gemini-cli) ## Setup the MCP Server Connect your AI client to Omnistrate MCP and authorize access to manage your Omnistrate infrastructure. ### Claude Code ``` # Install Claude Code npm install -g @anthropic-ai/claude-code # Configure Omnistrate MCP Server claude mcp add --scope user omnistrate omctl mcp start # Start coding with Claude claude ``` ### Claude.ai and Claude for Desktop Custom connectors using remote MCP are available on Claude and Claude Desktop for users on Pro, Max, Team, and Enterprise plans. 1. Open Settings in the sidebar 1. Navigate to Connectors and select Add custom connector 1. Configure the connector: 1. Name: `Omnistrate` 1. Command: `omnistrate-ctl mcp start` ### Cursor Add the snippet below to your project-specific or global `.cursor/mcp.json` file: ``` { "mcpServers": { "omnistrate": { "command": "omnistrate-ctl", "args": ["mcp", "start"] } } } ``` Once the server is added, Cursor will attempt to connect. Ensure your Omnistrate credentials are configured. ### VS Code with Copilot #### Installation Enable the mcp server with one command: ``` omnistrate-ctl mcp vscode enable ``` or manually: 1. Open the Command Palette (`Ctrl+Shift+P` on Windows/Linux or `Cmd+Shift+P` on macOS) 1. Run MCP: Add Server 1. Select Stdio 1. Enter the following details: 1. Command: `omnistrate-ctl` 1. Arguments: `mcp start` 1. Name: `Omnistrate` 1. Select Global or Workspace depending on your needs 1. Click Add #### Authorization Now that you've added Omnistrate MCP, let's start the server: 1. Open the Command Palette (`Ctrl+Shift+P` on Windows/Linux or `Cmd+Shift+P` on macOS) 1. Run MCP: List Servers 1. Select Omnistrate 1. Click Start Server 1. Ensure your Omnistrate credentials are configured in your environment ### Cline Add Omnistrate MCP to your Cline configuration: 1. Open Cline settings 1. Add a new MCP server: 1. Name: `Omnistrate` 1. Command: `omnistrate-ctl` 1. Arguments: `mcp start` 1. Save and restart Cline ### Windsurf Add the snippet below to your `mcp_config.json` file: ``` { "mcpServers": { "omnistrate": { "command": "omnistrate-ctl", "args": ["mcp", "start"] } } } ``` ### Gemini Code Assist To set up Omnistrate MCP with Gemini Code Assist: 1. Ensure you have Gemini Code Assist installed in your IDE 1. Add the following configuration to your `~/.gemini/settings.json` file: ``` { "mcpServers": { "omnistrate": { "command": "omnistrate-ctl", "args": ["mcp", "start"] } } } ``` 1. Restart your IDE to apply the configuration 1. Ensure your Omnistrate credentials are configured ### Gemini CLI Gemini CLI shares the same configuration as Gemini Code Assist. To set up Omnistrate MCP with Gemini CLI: 1. Ensure you have the Gemini CLI installed 1. Add the following configuration to your `~/.gemini/settings.json` file: ``` { "mcpServers": { "omnistrate": { "command": "omnistrate-ctl", "args": ["mcp", "start"] } } } ``` 1. Run the Gemini CLI and use the `/mcp list` command to see available MCP servers 1. Ensure your Omnistrate credentials are configured Setup steps may vary based on your MCP client version. Always check your client's documentation for the latest instructions. ## Agent instructions For optimal performance and safety when using Omnistrate MCP with AI agents, we recommend using the instructions provided in the [Omnistrate Agent Instructions](https://github.com/omnistrate-oss/agent-instructions) repository. These instructions are designed to: - Guide AI agents on best practices for interacting with Omnistrate - Ensure proper handling of sensitive operations - Improve the quality and reliability of agent-assisted workflows - Provide context-specific guidance for common Omnistrate tasks Visit the [agent-instructions repository](https://github.com/omnistrate-oss/agent-instructions) to access the recommended instructions and integrate them into your AI agent configuration. ## Using Skills ### Setup skills with CTL Setup the skills with one command ``` omnistrate-ctl agent init ``` ### For Claude Code Users Claude Code automatically discovers and uses skills defined in [CLAUDE.md](https://github.com/omnistrate-oss/agent-instructions/blob/main/CLAUDE.md). Simply start working with Omnistrate and Claude will invoke the appropriate skill based on your intent. ### For Other Agent Users 1. Read [AGENTS.md](https://github.com/omnistrate-oss/agent-instructions/blob/main/AGENTS.md) to understand available skills 1. Browse `skills/*/SKILL.md` files for workflow guidance 1. Consult `skills/*/*_REFERENCE.md` files for detailed syntax and examples 1. Configure your agent with the Omnistrate MCP server ### Quick Start **Designing a new SaaS application:** 1. Use the **omnistrate-sa** skill 1. Start with requirements (domain, scale, compliance, SLA) 1. Select appropriate technologies 1. Iteratively develop Docker Compose spec 1. Handoff vanilla compose to FDE skill **Onboarding a service:** 1. Use the **omnistrate-fde** skill 1. Start with your Docker Compose file (vanilla or from SA skill) 1. Transform to Omnistrate-native (ZERO parameterization initially) 1. Build, deploy, and iterate until RUNNING 1. Add API parameters incrementally when user requests customization **Debugging a deployment:** 1. Use the **omnistrate-sre** skill 1. Start with deployment status analysis 1. Follow the progressive debugging workflow 1. Identify root cause and resolve issues ## MCP Tools Required Both skills require the Omnistrate MCP server providing: - `mcp__ctl__account_*` - Cloud account management - `mcp__ctl__docs_*` - Documentation search - `mcp__ctl__build_compose` - Service builds - `mcp__ctl__service_plan_*` - Plan management - `mcp__ctl__instance_*` - Instance operations - `mcp__ctl__workflow_*` - Workflow analysis - `mcp__ctl__deployment-cell_*` - Kubernetes access ## Security best practices The MCP ecosystem and technology are evolving quickly. Here are our current best practices to help you keep your workspace secure: ### Verify your setup - Ensure `omnistrate-ctl` is installed from the official Omnistrate repository - Verify that your Omnistrate credentials are properly configured ### Trust and verification - Connecting to Omnistrate MCP grants the AI system you're using the same access as your Omnistrate user account - Review the permissions and access levels of each AI tool you connect ### Enable human confirmation - Always enable human confirmation in your workflows to maintain control and prevent unauthorized changes - Prevents accidental or harmful changes to your services and deployments ## Contributing We welcome contributions to improve the Omnistrate agent instructions! If you have suggestions for better workflows, additional skills, or improvements to existing guidance, please visit the [agent-instructions repository](https://github.com/omnistrate-oss/agent-instructions) and submit a pull request or open an issue. # Getting Started with Omnistrate This guide shows how to use Omnistrate to generate a private control plane for your SaaS product. Follow these steps to define your application, connect the infrastructure it needs, and create a control plane that can provision, operate, meter, and manage customer deployments across your supported deployment models. ## What makes up your SaaS Product? Each SaaS product built with Omnistrate is a collection of software, resources, parameters, features, and infrastructure configurations. Omnistrate uses your product definition to generate the APIs, workflows, customer portal, deployment automation, observability, and operational controls that make up your control plane. Omnistrate provides different options for specifying your SaaS product. The following steps guide you through those options. ## Step 1: Sign up to Omnistrate Before building your SaaS Product, you need to set up your Omnistrate account: 1. **Sign up for Omnistrate** at [omnistrate.cloud](https://omnistrate.cloud) using your email, Google account, or GitHub account. 1. When signing up with email **Complete account verification** through the email confirmation process Note You will receive a confirmation email. Please check your spam folder if you do not see the email in your inbox. Once your account is set up, you'll have access to the Omnistrate dashboard where you can manage your SaaS Products, environments, and deployments. ## Step 2: Onboard Your Account After signing up, you need to onboard your cloud account. Omnistrate uses your cloud account (such as AWS, GCP, or Azure) to provision and manage the infrastructure required for your SaaS Product. This model ensures that your application and customer data remain within your own secure cloud environment. Onboarding involves granting Omnistrate the necessary permissions to automate deployments and manage resources on your behalf. For detailed instructions, follow the [Account Onboarding Guide](https://docs.omnistrate.com/getting-started/account-onboarding/index.md). ## Step 3: Install Omnistrate CLI (CTL) The Omnistrate Command Line Interface (CTL) is essential for building and managing your SaaS Products: 1. **Download and install** the CTL from [ctl.omnistrate.cloud/install](https://ctl.omnistrate.cloud/install/) 1. **Verify installation** by running: ``` omnistrate-ctl --help ``` 1. **Login to your account**: ``` omnistrate-ctl login ``` The CTL provides commands for building services, managing deployments, and automating CI/CD workflows. ## Step 4: Onboard your application Omnistrate provides flexible options to define your application and generate its control plane from the development approach you already use. Choose the model that best fits your current situation: ### Start From your Application If you're building a new service or want to containerize your existing application, these options are ideal: #### a) Build from Source Repository First, configure your MCP Server [here](https://docs.omnistrate.com/getting-started/mcp-server/index.md) And then, build directly from your source code repository. Omnistrate will handle the containerization and deployment process. - **Best for**: Applications with source code in Git repositories - **Benefits**: Automated builds from code changes, simplified CI/CD integration - **Guide**: [Build from Repository](https://docs.omnistrate.com/getting-started/build-from-repo/index.md) #### b) Build from Container Image Use your existing Docker container images to define your service. - **Best for**: Applications already containerized or with existing Docker workflows - **Benefits**: Quick setup, leverages existing container investments - **Guide**: [Getting Started with Compose](https://docs.omnistrate.com/getting-started/build-from-repo/index.md) If you need help iterating with your compose spec, you can configure your MCP Server [here](https://docs.omnistrate.com/getting-started/mcp-server/index.md) #### c) Build from Docker Compose Transform your Docker Compose configuration into a multi-tenant SaaS Product. - **Best for**: Applications defined with Docker Compose files - **Benefits**: Familiar format, quick migration from development to production - **Guide**: [Getting Started with Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) ##### Setup your AI Agent Enhance your development workflow by connecting AI tools to Omnistrate using the [Model Context Protocol (MCP)](https://modelcontextprotocol.io). This allows AI assistants to interact directly with your Omnistrate services and infrastructure. With Omnistrate MCP, AI assistants can help you: - Design and architect a new service plan - Transform Docker Compose files into SaaS offerings - Debug deployments and troubleshoot issues For VS Code: ``` omnistrate-ctl mcp vscode enable omnistrate-ctl agent init ``` For Claude Code: ``` omnistrate-ctl mcp claude enable omnistrate-ctl agent init ``` For detailed setup instructions for all supported AI tools, see the [MCP Server Configuration Guide](https://docs.omnistrate.com/getting-started/mcp-server/index.md). This step is optional and recommended. ### Start from Existing Infrastructure Packages If you already have infrastructure automation in place, these options preserve your existing investments: #### d) Build from Helm Charts Integrate your existing Helm charts with Omnistrate's multi-tenancy and automation capabilities. - **Best for**: Kubernetes-native applications with existing Helm charts - **Benefits**: Preserves existing Kubernetes configurations and deployment patterns - **Guide**: [Helm Charts Integration](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) #### e) Build from Terraform Leverage your existing Terraform modules to define infrastructure and application components. - **Best for**: Infrastructure-heavy applications or multi-cloud deployments - **Benefits**: Reuses existing infrastructure code, supports complex architectures - **Guide**: [Terraform Integration](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md) #### f) Build from Kubernetes Operators Use your Kubernetes Operators to manage application lifecycle on Omnistrate, including operator CRDs, system workflows, backup/restore, scaling, start/stop, and provider-defined custom operations. - **Best for**: Applications with complex operational requirements or custom controllers - **Benefits**: Preserves operator logic, adds Customer Portal, APIs, tenant management, fleet operations, and workflow-driven day-2 automation - **Guide**: [Kubernetes Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md) #### g) Build from Kustomize Transform your Kustomize configurations into managed SaaS Products. - **Best for**: Teams using Kustomize for Kubernetes configuration management - **Benefits**: Maintains configuration layering and customization patterns - **Guide**: [Kustomize Integration](https://docs.omnistrate.com/getting-started/build-from-kustomize/index.md) ## Next Steps After completing these steps, you'll have: - ✅ An active Omnistrate account - ✅ A defined service specification matching your infrastructure approach - ✅ The CTL installed and configured - ✅ Your first private control plane generated with Omnistrate ### Continue Your Journey 1. **Configure Environments**: Set up development, staging, and production environments 1. **Set Up Customer Portal**: Enable customer self-service with the [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) 1. **Implement CI/CD**: Automate deployments with [CI/CD workflows](https://github.com/omnistrate-community/ci-cd-example) 1. **Add Advanced Features**: Explore [integrations](https://docs.omnistrate.com/build-guides/integrations/index.md), [monitoring](https://docs.omnistrate.com/build-guides/service-visibility/index.md), and [billing](https://github.com/omnistrate-community/usage-export-clazar-recipe) For API-based automation and advanced control, see the [API documentation](https://api.omnistrate.cloud/docs/external/). For account management, including offboarding procedures, see [Account Offboarding](https://docs.omnistrate.com/getting-started/account-onboarding/#offboarding). For any questions, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com). We would love to understand your use-case and assist you. # Zero to SaaS Product in 5 minutes ## From Source Code to SaaS Product This video demonstrates how Omnistrate enables rapid deployment of a PG vector database as a service directly from a Git repository, using a streamlined CLI workflow. ### Key Features to Build your SaaS Product - **Single Command Deployment**: Omnistrate’s CLI tool allows seamless conversion of a Git repository into a fully operational RDS-style PG vector database service with one command. - **Deployment Models**: - *Hosted SaaS*: The service provider manages the service in its own account. - *BYOC*: Customers deploy database instances in their own AWS, GCP, or Azure cloud accounts for maximum control. - *Air-gapped*: Customers deploy the service in isolated environments without a live control-plane connection. ## Auto-generated Customer Portal The video demonstrates how Omnistrate enables service providers to offer a seamless customer experience through a branded customer portal that supports single sign-on and secure access. ### Key Features for Customer Portal - **Cloud Account Onboarding**: Customers can quickly onboard their cloud accounts and deploy services like PG vector instances within minutes. - **Deployment Management**: The portal streamlines deployment creation, monitoring, and management, showing real-time load and health status, configurations, connectivity endpoints, and nodes list. - **Transparency & Monitoring**: Customers access audit trails, detailed metrics, and logs to facilitate ongoing monitoring and troubleshooting. - **Resilience & Recovery**: Full control over key management actions (reboot, start, stop, capacity adjustment) and disaster recovery features, including point-in-time recovery, ensure enhanced data integrity. - **Dashboard & Collaboration**: A dashboard provides a high-level view; customers can review audit logs, receive system notifications, invite team members, and assign role-based access. - **Subscription & Governance**: Integrated tools allow subscription management, governance over deployments, and real-time usage monitoring. - **API & CLI Generation**: Omnistrate automates the creation of REST API and CLI tools for advanced tasks and automation. ## Day-2 Overview The video highlights Omnistrate's comprehensive management tools, enabling operators to efficiently oversee and maintain cloud services. ### Key Features for Operations - **Fleet Dashboard**: Provides a high-level overview of all managed services, enabling operators to quickly search and locate specific services, users, or instances. - **Workflow Tracking & Debug Events**: Monitors every instance operation by tracking workflow progress and surfacing real-time debug events, giving full visibility into status and potential issues. - **Resource Monitoring**: Operators have access to audit logs, performance metrics, and detailed logs for each resource instance, supporting effective monitoring and troubleshooting. - **Service Plans & Version Rollouts**: All service plan versions are accessible, with automated, scheduled version rollouts across instances, ensuring seamless upgrades for customers. - **Tenant Management & Security**: Simplifies managing tenant access and settings, including subscriptions and instances, and ensures secure and isolated data environments. - **Advanced Operations**: Supports scalable operations with features like billing, intelligent monitoring, backups, compliance, CI/CD integration, and notifications. ## Dive Deeper Ready to build your own SaaS Product? Start with the deployment strategy that matches your application: - [Build from Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) — Ideal if you have a Docker Compose setup - [Build from Helm](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) — Best for Kubernetes-native applications with Helm charts - [Build from Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md) — Leverage Kubernetes Operators for stateful workloads - [Build from Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md) — Use Terraform for cloud infrastructure provisioning - [Build from Kustomize](https://docs.omnistrate.com/getting-started/build-from-kustomize/index.md) — Deploy with Kustomize overlays for Kubernetes - [Build from Source Repo](https://docs.omnistrate.com/getting-started/build-from-repo/index.md) — Build directly from your Dockerfile and source code Then explore the [Build Guide](https://docs.omnistrate.com/build-guides/overview/index.md) for core concepts like deployment models, tenancy types, API parameters, and more. # AWS account onboarding with CTL Use this guide when your hosted plan or BYOA plan deploys to AWS. ## Region model AWS account configs do not use bindings. Once the account is verified and `READY`, the same account config can be used across supported AWS regions. Pick the target region at deployment time with `--region`. ## Prerequisites - An AWS account ID. - Permission to run the Omnistrate onboarding bootstrap in that AWS account. - `omnistrate-ctl` installed and logged in. ## Hosted flow ### 1. Register the hosted account config For AWS, use `--skip-wait` because the account will not become `READY` until you complete the bootstrap step in AWS. ``` omnistrate-ctl account create aws-hosted \ --aws-account-id 123456789012 \ --skip-wait ``` ### 2. Open the bootstrap action ``` omnistrate-ctl account describe aws-hosted ``` In the account TUI: - Select **Actions -> Bootstrap** and press `o` to open the CloudFormation URL. - If you need the alternate flow without the load-balancer policy, select **Bootstrap (No LB)**. - Press `c` to copy the selected URL instead of opening it. ### 3. Wait for the account to become ready After the CloudFormation or Terraform bootstrap completes, rerun: ``` omnistrate-ctl account describe aws-hosted ``` Continue until the account status is `READY`. ### 4. Create the first Hosted SaaS deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider aws \ --region us-east-1 \ --param-file ./params.json ``` ## BYOA flow ### 1. Create the customer onboarding instance Use `--skip-wait` here for the same reason: the customer-facing bootstrap must run in AWS before the backing account can become `READY`. ``` omnistrate-ctl account customer create \ --service \ --environment \ --plan \ --aws-account-id 123456789012 \ --skip-wait ``` In production environments, the command uses the calling user's subscription by default. Add `--subscription-id` or `--customer-email` only if you need to onboard on behalf of a different production subscription. ### 2. Find the backing account config and complete bootstrap ``` omnistrate-ctl account customer describe -o json ``` Copy `summary.accountConfigID`, then open the backing account config: ``` omnistrate-ctl account describe ``` Use the **Actions -> Bootstrap** entry in the hosted account TUI to open or copy the AWS bootstrap URL, then complete the bootstrap in the customer AWS account. ### 3. Verify the customer onboarding instance ``` omnistrate-ctl account customer describe omnistrate-ctl account customer list --cloud-provider aws ``` Continue until the onboarding instance and the backing account are both ready to use. ### 4. Create the BYOA deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider aws \ --region us-east-1 \ --customer-account-id \ --param-file ./params.json ``` ## Update and delete AWS does not require account-level updates to add regions. To reuse the same account in another AWS region, deploy with a different `--region`. Useful lifecycle commands: ``` # Inspect all customer AWS onboarding instances omnistrate-ctl account customer list --cloud-provider aws # Delete a customer onboarding instance omnistrate-ctl account customer delete ``` # Azure account onboarding with CTL Use this guide when your hosted plan or BYOA plan deploys to Azure. ## Region model Azure account configs do not use bindings. Once the account is verified and `READY`, the same account config can be reused across supported Azure regions. Pick the target region during deployment creation. ## Prerequisites - An Azure subscription ID and tenant ID. - Permission to run the Omnistrate bootstrap command in Azure Cloud Shell. - The Azure and Microsoft Entra roles called out in [Account Onboarding](https://docs.omnistrate.com/getting-started/account-onboarding/#azure). - `omnistrate-ctl` installed and logged in. ## Hosted flow ### 1. Register the hosted account config For Azure, use `--skip-wait` because the account cannot become `READY` until the bootstrap command is executed in Azure Cloud Shell. ``` omnistrate-ctl account create azure-hosted \ --azure-subscription-id 00000000-0000-0000-0000-000000000000 \ --azure-tenant-id 11111111-1111-1111-1111-111111111111 \ --skip-wait ``` ### 2. Copy the bootstrap shell command ``` omnistrate-ctl account describe azure-hosted ``` In the account TUI: - Select **Actions -> Bootstrap**. - Press `c` to copy the bootstrap shell command. - Paste that command into Azure Cloud Shell and run it. ### 3. Wait for the account to become ready After the bootstrap command succeeds, rerun: ``` omnistrate-ctl account describe azure-hosted ``` Continue until the account status is `READY`. ### 4. Create the first Hosted SaaS deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider azure \ --region eastus \ --param-file ./params.json ``` ## BYOA flow ### 1. Create the customer onboarding instance ``` omnistrate-ctl account customer create \ --service \ --environment \ --plan \ --azure-subscription-id 00000000-0000-0000-0000-000000000000 \ --azure-tenant-id 11111111-1111-1111-1111-111111111111 \ --skip-wait ``` In production environments, the command uses the calling user's subscription by default. Add `--subscription-id` or `--customer-email` only if you need to onboard on behalf of a different production subscription. ### 2. Find the backing account config and run the bootstrap command ``` omnistrate-ctl account customer describe -o json ``` Copy `summary.accountConfigID`, then open the backing account config: ``` omnistrate-ctl account describe ``` Select **Actions -> Bootstrap**, copy the Azure Cloud Shell command, and run it in the target customer subscription. ### 3. Verify the customer onboarding instance ``` omnistrate-ctl account customer describe omnistrate-ctl account customer list --cloud-provider azure ``` ### 4. Create the BYOA deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider azure \ --region eastus \ --customer-account-id \ --param-file ./params.json ``` ## Update and delete Azure does not require account-level updates to add regions. Reuse the same account config and choose a different Azure region when you create the instance. Useful lifecycle commands: ``` # Inspect all customer Azure onboarding instances omnistrate-ctl account customer list --cloud-provider azure # Delete a customer onboarding instance omnistrate-ctl account customer delete ``` # GCP account onboarding with CTL Use this guide when your hosted plan or BYOA plan deploys to Google Cloud. ## Region model GCP account configs do not use bindings. Once the account is verified and `READY`, the same account config can be reused across supported GCP regions. Pick the target region during deployment creation. ## Prerequisites - A GCP project ID and project number. - Permission to run the Omnistrate bootstrap command in Google Cloud Shell for that project. - `omnistrate-ctl` installed and logged in. ## Hosted flow ### 1. Register the hosted account config For GCP, use `--skip-wait` because the account cannot become `READY` until the bootstrap command is executed in Google Cloud Shell. ``` omnistrate-ctl account create gcp-hosted \ --gcp-project-id my-project \ --gcp-project-number 123456789012 \ --skip-wait ``` ### 2. Copy the bootstrap shell command ``` omnistrate-ctl account describe gcp-hosted ``` In the account TUI: - Select **Actions -> Bootstrap**. - Press `c` to copy the bootstrap shell command. - Paste that command into Google Cloud Shell and run it. ### 3. Wait for the account to become ready After the bootstrap command succeeds, rerun: ``` omnistrate-ctl account describe gcp-hosted ``` Continue until the account status is `READY`. ### 4. Create the first Hosted SaaS deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider gcp \ --region us-central1 \ --param-file ./params.json ``` ## BYOA flow ### 1. Create the customer onboarding instance ``` omnistrate-ctl account customer create \ --service \ --environment \ --plan \ --gcp-project-id my-project \ --gcp-project-number 123456789012 \ --skip-wait ``` In production environments, the command uses the calling user's subscription by default. Add `--subscription-id` or `--customer-email` only if you need to onboard on behalf of a different production subscription. ### 2. Find the backing account config and run the bootstrap command ``` omnistrate-ctl account customer describe -o json ``` Copy `summary.accountConfigID`, then open the backing account config: ``` omnistrate-ctl account describe ``` Select **Actions -> Bootstrap**, copy the Cloud Shell command, and run it in the target customer project. ### 3. Verify the customer onboarding instance ``` omnistrate-ctl account customer describe omnistrate-ctl account customer list --cloud-provider gcp ``` ### 4. Create the BYOA deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider gcp \ --region us-central1 \ --customer-account-id \ --param-file ./params.json ``` ## Update and delete GCP does not require account-level updates to add regions. Reuse the same account config and choose a different deployment region when you create the instance. Useful lifecycle commands: ``` # Inspect all customer GCP onboarding instances omnistrate-ctl account customer list --cloud-provider gcp # Delete a customer onboarding instance omnistrate-ctl account customer delete ``` # Use Nebius vLLM with Opencode Use this example if you want to publish a Nebius-hosted vLLM service plan with `omctl`, create a GPU-backed instance, expose the model over the Omnistrate-managed DNS endpoint, and use that endpoint from Opencode. This example builds on the account setup in [Nebius account onboarding with CTL](https://docs.omnistrate.com/getting-started/onboarding/nebius/index.md). - **model**: `Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled` - **served model alias**: `qwen3.5-27b-claude-4.6-opus-reasoning-distilled` - **tokenizer override**: `Qwen/Qwen3.5-27B` - **public OpenAI-compatible vLLM API** over HTTPS on the Omnistrate-managed DNS endpoint - **tool-calling enabled** so Opencode can use `/v1/chat/completions` - **Omnistrate dashboards** for both vLLM metrics and NVIDIA GPU metrics ## 1. Choose the spec Clone the example from . The repository contains two variants: - `spec.yaml`: single-node Nebius deployment on a single H200 GPU - `spec-gpu-cluster.yaml`: multi-GPU Nebius GPU-cluster deployment on H200 with tensor parallelism enabled Use `spec.yaml` if you want the simplest path for Opencode. Use `spec-gpu-cluster.yaml` only if you already have a Nebius GPU cluster and want to run the same model with 8 GPUs. Before building that spec, replace the placeholder `GpuClusterID: compute-cluster-id` with your real Nebius GPU cluster ID. ## 2. Build and release the service plan Run the build from the directory that contains the spec you want to publish. Single-GPU example: ``` omctl build -f spec.yaml \ --product-name 'Nebius' \ --spec-type ServicePlanSpec \ --release-as-preferred ``` GPU-cluster example: ``` omctl build -f spec-gpu-cluster.yaml \ --product-name 'Nebius' \ --spec-type ServicePlanSpec \ --release-as-preferred ╭─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ environment plan_id plan_name release_description service_id service_name version version_set_status │ │─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ │ Dev pt-VRWPNLImF0 Nebius vLLM GPU Inference Initial Release Version Set s-GuWdEcfcP5 Nebius vLLM Demo 1.0 Preferred │ ╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ Check the service plan result at: https://omnistrate.cloud/product-tier?serviceId=s-GuWdEcfcP5&environmentId=se-BGBkF9Nwqq Access your SaaS offer at: https://saasportal.instance-w6vidhd14.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-GuWdEcfcP5&environmentId=se-BGBkF9Nwqq ``` This creates or updates the service named `Nebius` and releases the plan version as the preferred version in the target environment. In this example, the plan name inside the spec is `Nebius vLLM GPU Inference`. ## 3. Create a Nebius vLLM instance Once the plan is released, create an instance in a region backed by a `READY` Nebius binding. ``` omctl instance create \ --service=Nebius \ --environment=dev \ --plan='Nebius vLLM GPU Inference' \ --resource=vllm \ --cloud-provider=nebius \ --region= ╭────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ cloud_provider environment instance_id plan region resource service status subscription_id tags version │ │────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ │ nebius Dev Nebius vLLM GPU Inference us-central1 vllm Nebius vLLM Demo DEPLOYING sub-6OR7umU3Ei 1.0 │ ╰────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ ``` If you are using a BYOA plan instead of a hosted plan, add `--customer-account-id ` to the create command. Keep the returned instance ID. You will use it to inspect the endpoint, configure Opencode, and later delete the deployment. If this is the first deployment in that Nebius tenant and region, expect the initial infrastructure bring-up to take longer than a normal redeploy. You can track the instance status with: ``` omctl instance describe # Filter using jq omctl instance describe | jq '.consumptionResourceInstanceResult.status' "RUNNING" ``` If something is not progressing, inspect the deployment with: ``` omctl instance debug ``` ## 4. Get the generated inference endpoint List the instance endpoints: ``` omctl instance list-endpoints ╭───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ endpoint_name endpoint_type network_type ports resource_name status url │ │───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ │ api additional PUBLIC 443 vllm HEALTHY r-o3davudsia..hc-xyxgpfjol.us-central1.nebius.f2e0a955bb84.cloud │ ╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ ``` You should see a public `api` endpoint pointing at the Omnistrate-managed DNS name for the vLLM service. ## 5. Verify the endpoint before wiring in Opencode Health check: ``` curl -i https:///health ``` List the served models: ``` curl https:///v1/models ``` You should see the served model ID: ``` qwen3.5-27b-claude-4.6-opus-reasoning-distilled ``` Run a basic chat completion: ``` curl -sS \ -H 'Content-Type: application/json' \ -d '{ "model": "qwen3.5-27b-claude-4.6-opus-reasoning-distilled", "messages": [ { "role": "user", "content": "Write a Python function that reverses a linked list." } ], "max_tokens": 512, "temperature": 0 }' \ https:///v1/chat/completions ``` Because Opencode uses tool-enabled OpenAI-compatible chat requests, it is worth verifying tool calling too: ``` curl -sS \ -H 'Content-Type: application/json' \ -d '{ "model": "qwen3.5-27b-claude-4.6-opus-reasoning-distilled", "messages": [ { "role": "user", "content": "What is 2 + 2? Use the tool if needed." } ], "tools": [ { "type": "function", "function": { "name": "echo", "description": "Echo the provided text", "parameters": { "type": "object", "properties": { "text": { "type": "string" } }, "required": ["text"] } } } ], "tool_choice": "auto", "max_tokens": 256 }' \ https:///v1/chat/completions ``` The shipped spec already enables the vLLM flags required for this flow, including `--enable-auto-tool-choice`, `--tool-call-parser qwen3_coder`, and `--reasoning-parser qwen3`. If this request returns a `400` complaining that `auto` tool choice is not enabled, the instance is still running an older plan version or an older chart release. ## 6. Configure Opencode Point Opencode at the Nebius vLLM endpoint by editing `~/.config/opencode/opencode.json`. ``` { "$schema": "https://opencode.ai/config.json", "provider": { "nebius-vllm": { "npm": "@ai-sdk/openai-compatible", "name": "Nebius vLLM Qwen 3.5 Distilled", "options": { "baseURL": "https:///v1" }, "models": { "qwen3.5-27b-claude-4.6-opus-reasoning-distilled": { "name": "qwen3.5-27b-claude-4.6-opus-reasoning-distilled" } } } } } ``` Notes: - Replace `` with the DNS name returned by `omctl instance list-endpoints`. - Do not append `:443` or `:8000`; the ingress already exposes the service over HTTPS on the default port. - If your endpoint requires authentication, add `options.apiKey` or `options.headers`. - If you want Opencode to understand model limits more precisely, add a `limit` block under the model entry. After saving the config, start Opencode and run: ``` /models ``` Select: - provider: `nebius-vllm` - model: `qwen3.5-27b-claude-4.6-opus-reasoning-distilled` At that point Opencode will use your Nebius-hosted vLLM endpoint for chat and tool-enabled coding requests. ## 7. Inspect the Omnistrate dashboards This example includes Omnistrate dashboard integrations for both vLLM and NVIDIA GPU telemetry. Open the dashboards with: ``` omctl instance dashboard ``` Out of the box, the dashboards expose: - vLLM request concurrency, throughput, latency, KV cache usage, and prefix cache hit rate - NVIDIA GPU inventory, utilization, framebuffer usage, PCIe throughput, thermals, power, and error signals ## 8. Delete the instance Delete the deployment when you are finished: ``` omctl instance delete ``` Add `-y` if you want to skip the confirmation prompt: ``` omctl instance delete -y ``` # Nebius account onboarding with CTL Use this guide when your hosted plan or BYOA plan deploys to Nebius. ## Why Nebius is different Nebius account configs are **tenant-scoped**. Instead of a single account-wide identity, Omnistrate expects one or more **bindings** inside the account config. A binding is the combination of: - a Nebius project ID, - a Nebius service account ID, - a Nebius public key ID, and - the matching private key PEM. Each binding enables a specific region in your Nebius tenant for deployment after Omnistrate verifies it and confirms the service-account key belongs to the configured tenant. In practice this means: - use **one binding per Nebius project/region**, - keep **all existing bindings** in the file when adding a new region, and - only deploy into regions whose bindings show `READY` in `account describe`. One binding also acts as the **artifact bucket owner** for the account config. Omnistrate uses that binding's project for Nebius deployment artifacts (Helm charts, Terraform files, etc.). When an account config contains more than one binding, exactly one binding must own the artifact bucket. ## Prerequisites - A Nebius tenant. - One Nebius project for each region you want to enable. - A human operator with enough Nebius IAM access to create service accounts and service-account authorized keys. - `omnistrate-ctl` installed and logged in. Useful Nebius references: - [Nebius service accounts](https://docs.nebius.com/iam/service-accounts/manage) - [Nebius service-account authentication and authorized keys](https://docs.nebius.com/iam/service-accounts/authentication) ## Create the Nebius service account and key Repeat this process for every Nebius project that you want to enable as a binding. ### 1. Create or choose the target project Create one Nebius project per deployment region you want Omnistrate to use. Keep the project ID for the binding file. ### 2. Create a service account in that project In Nebius Console: 1. Open the target project. 1. Go to **IAM** or **Service Accounts**. 1. Create a service account such as `omnistrate-eu-north1`. 1. Save the **service account ID**. ### 3. Grant the service account project permissions Add the service account to a project-level group or role set that can create and manage the infrastructure Omnistrate needs in that project. A project-scoped admin-style role is required to allow Omnistrate to manage your infrastructure. ### 4. Create an authorized key for the service account Upload a public key and save the matching private key PEM for the binding file. Save these values for the binding: - **project ID** - **service account ID** - **public key ID** - **private key PEM** Warning Treat the private key PEM like any other long-lived secret. Store it outside source control. The recommended pattern in this guide is to reference a local file from the bindings YAML. ## Artifact bucket owner behavior For Nebius, Omnistrate stores deployment artifacts in the project owned by one binding. - On the **first** account onboarding, if the bindings file does not specify an owner, Omnistrate picks the **first binding in the file** as the artifact-bucket owner. - On later full-replacement binding updates, Omnistrate preserves the existing artifact-bucket owner as long as that binding stays in the set. - If you remove the current artifact-owner binding from a later update, the update is rejected because Omnistrate would no longer know which Nebius project should keep artifact ownership. - In practice, put your most stable Nebius project first when you onboard the account for the first time. ## Create the Nebius bindings file A bindings file can carry one or more Nebius bindings. Omnistrate resolves the region during verification, so you do **not** specify a region in the YAML. ``` bindings: - projectID: project-e00project serviceAccountID: serviceaccount-e00serviceacct publicKeyID: publickey-e00publickey privateKeyPEMFile: /path/to/private-key.pem - projectID: project-e00anotherproject serviceAccountID: serviceaccount-e00anotherserviceacct publicKeyID: publickey-e00anotherpublickey privateKeyPEMFile: /path/to/another-private-key.pem ``` Notes: - `privateKeyPEMFile` is recommended. - `privateKeyPEM` also works if you must inline the key material. - On first onboarding, Omnistrate chooses the **first binding in the file** as the artifact-bucket owner. - Keep the file as the **full desired binding set**. Updates replace the complete set, not just one entry. - Keep the originally selected artifact-owner binding in later replacement files so artifact storage stays attached to the same Nebius project. If your plan also provisions Nebius resources through Terraform, see [Nebius Terraform authentication](https://docs.omnistrate.com/getting-started/build-from-terraform/#nebius-terraform-authentication). ## Hosted flow ### 1. Create the hosted account config ``` omnistrate-ctl account create nebius-hosted \ --nebius-tenant-id tenant-id \ --nebius-bindings-file ./nebius-bindings.yaml ``` Because Nebius verification is direct, you can let the command wait for `READY`. Add `--skip-wait` only if you want to return immediately. ### 2. Inspect binding verification ``` omnistrate-ctl account describe nebius-hosted omnistrate-ctl account describe nebius-hosted -o json ``` Only use regions whose bindings are `READY`. ### 3. Create the first Hosted SaaS deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider nebius \ --region eu-north1 \ --param-file ./params.json ``` The `--region` value must match one of the ready binding regions surfaced by `account describe`. ## BYOA flow ### 1. Create the customer onboarding instance ``` omnistrate-ctl account customer create \ --service \ --environment \ --plan \ --nebius-tenant-id tenant-e00ezh17k22wmwq5f0 \ --nebius-bindings-file ./nebius-bindings.yaml ``` The command uses the calling user's subscription by default. If you need to onboard on behalf of another production subscription, add either `--subscription-id ` or `--customer-email `. ### 2. Verify the customer onboarding instance ``` omnistrate-ctl account customer list --cloud-provider nebius omnistrate-ctl account customer describe omnistrate-ctl account customer describe -o json ``` Use the JSON output if you want the backing `accountConfigID` and the raw Nebius binding verification details. ### 3. Create the BYOA deployment ``` omnistrate-ctl instance create \ --service \ --environment \ --plan \ --resource \ --cloud-provider nebius \ --region eu-north1 \ --customer-account-id \ --param-file ./params.json ``` For BYOA plans, `--customer-account-id` is the onboarding instance ID returned by `account customer create`. ## Author Nebius plans ### Nebius compute options ``` services: - name: vllm compute: instanceTypes: - name: 8gpu-128vcpu-1600gb cloudProvider: nebius platform: gpu-h100-sxm configurationOverrides: GpuClusterID: computegpucluster-e00abcxfadyhcvpkdx ``` Notes: - The instance type `name` is the Nebius preset. If you use `apiParam` instead, that parameter must resolve to a Nebius preset at deploy time. - `platform` is required on Nebius instance-type entries. - `configurationOverrides.GpuClusterID` is optional and Nebius-only. Use it when the preset must land in a specific Nebius GPU cluster. - Nebius GPU cluster placement is validated against the selected `platform` and preset. ### Nebius shared file systems Nebius plans can also use Omnistrate-managed shared file systems. In compose-based plans: - define the shared volume with `driver: sharedFileSystem`, - put the Nebius backend config on the root volume `driver_opts`, - and reference the matching Nebius `clusterStorageType` in each mounted Resource. Example: ``` volumes: shared_models: driver: sharedFileSystem driver_opts: nebiusFilesystemType: NETWORK_SSD nebiusFilesystemSizeGi: 1024 services: api: volumes: - source: shared_models target: /models type: volume x-omnistrate-storage: nebius: clusterStorageType: NEBIUS::FILESYSTEM_NETWORK_SSD ``` Supported Nebius backend mappings: - `NETWORK_SSD` -> `NEBIUS::FILESYSTEM_NETWORK_SSD` - `NETWORK_HDD` -> `NEBIUS::FILESYSTEM_NETWORK_HDD` - `WEKA` -> `NEBIUS::FILESYSTEM_WEKA` - `VAST` -> `NEBIUS::FILESYSTEM_VAST` Notes: - The filesystem size and backend type live on the shared volume `driver_opts`, not on the per-resource mount block. - Omnistrate rejects Nebius shared file systems for `OMNISTRATE_MULTI_TENANCY` plans. Use Nebius dedicated or custom tenancy instead. - See [Shared File System](https://docs.omnistrate.com/infra-guides/shared-file-system/index.md) and [Compose Spec](https://docs.omnistrate.com/build-guides/compose-spec/#volumesdriver-sharedfilesystem) for the full schema. ### Nebius Terraform authentication ``` terraformConfigurations: configurationPerCloudProvider: nebius: terraformPath: /terraform/nebius serviceAccountID: serviceaccount-e00vqdp9fskhmmaan8 publicKeyID: publickey-e00h9scsyy9mbefrjf privateKeyPEM: $secret.nebiusTerraformPrivateKey gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Notes: - Prefer `$secret.` for `privateKeyPEM`; Omnistrate resolves the secret before running Terraform. - When these fields are present, Omnistrate uses this service-account key for the Terraform resource instead of the default host-cluster Terraform identity. - See [Building your SaaS Product using Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/#nebius-terraform-authentication) and [Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/#terraform-per-cloud-provider-configuration-schema) for the full schema. ## Add a new region later To enable another Nebius region, create another project-level binding and replace the full bindings file. ### Hosted account update ``` omnistrate-ctl account update nebius-hosted \ --nebius-bindings-file ./nebius-bindings.yaml ``` ### BYOA backing account update ``` omnistrate-ctl account customer update \ --nebius-bindings-file ./nebius-bindings.yaml ``` After the update completes: 1. Rerun `account describe` or `account customer describe`. 1. Confirm the new binding reaches `READY`. 1. Start deploying into the newly discovered region. Warning Nebius binding updates are full replacements. If your current file has two working bindings and you add a third, keep all three in the replacement file. Warning The binding that was first in your original onboarding file becomes the artifact-bucket owner. Keep that binding in later replacement files unless you are intentionally migrating Nebius artifact storage. ## Examples For an end-to-end workload example that builds and publishes a service plan, creates a GPU-backed instance, tests the external endpoint, and deletes the deployment again, see [Deploy vLLM on Nebius](https://docs.omnistrate.com/getting-started/onboarding/nebius-vllm/index.md). ## Delete and offboard Useful lifecycle commands: ``` # List all Nebius customer onboarding instances omnistrate-ctl account customer list --cloud-provider nebius # Delete a customer onboarding instance omnistrate-ctl account customer delete ``` For the offboarding order and cleanup rules, follow [Account Offboarding](https://docs.omnistrate.com/getting-started/account-onboarding/#offboarding). # Cloud account guides for CTL These guides show how to onboard cloud accounts with `omnistrate-ctl` for both hosted plans and customer-managed BYOA plans. Note The examples below use `omnistrate-ctl`. If you use the shorter `omctl` alias locally, replace the command name directly. ## Hosted vs BYOA - **Hosted plans** run in an account that you operate. Start with `omnistrate-ctl account create`. - **BYOA plans** run in an account owned by the customer or by a customer-specific deployment boundary. Start with `omnistrate-ctl account customer create`. - **Deployments** use `omnistrate-ctl instance create`. BYOA plans additionally require `--customer-account-id`. ## Provider guide matrix | Provider | Hosted account flow | BYOA account flow | Region enablement model | Guide | | --------------- | ----------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------- | -------------------------------------------------------------------------------- | | AWS | `account create` + CloudFormation/Terraform bootstrap | `account customer create` + customer bootstrap | Region chosen at deployment time | [AWS](https://docs.omnistrate.com/getting-started/onboarding/aws/index.md) | | GCP | `account create` + Cloud Shell bootstrap | `account customer create` + customer bootstrap | Region chosen at deployment time | [GCP](https://docs.omnistrate.com/getting-started/onboarding/gcp/index.md) | | Azure | `account create` + Azure Cloud Shell bootstrap | `account customer create` + customer bootstrap | Region chosen at deployment time | [Azure](https://docs.omnistrate.com/getting-started/onboarding/azure/index.md) | | Nebius | `account create` + binding verification | `account customer create` + binding verification | Ready bindings determine eligible regions | [Nebius](https://docs.omnistrate.com/getting-started/onboarding/nebius/index.md) | | BYOC On-Premise | Not applicable | `account customer create` + `account customer install-kit` | Customer cluster selected during onboarding | [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) | ## Shared lifecycle commands Use these commands across all providers: ``` # Inspect a hosted account config omnistrate-ctl account describe # Rename an existing hosted account config or replace Nebius bindings omnistrate-ctl account update [flags] # List customer BYOA onboarding instances omnistrate-ctl account customer list [flags] # Inspect a customer BYOA onboarding instance omnistrate-ctl account customer describe # Update a customer BYOA onboarding instance or its backing account config omnistrate-ctl account customer update [flags] # Delete a customer BYOA onboarding instance omnistrate-ctl account customer delete ``` For production environments, `omnistrate-ctl account customer create` uses the calling user's subscription by default. Override that with `--subscription-id` or `--customer-email` only when you need to onboard on behalf of a different subscription. All account commands accept a `--tags "key=value,key2=value2"` flag to set custom tags on the account configuration — for example to mark environments or cost centers, and to steer per-account behavior such as [conditional amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/#conditional-amenities). See [Account Tags](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#account-tags) for the full semantics. ## Before you start - Install the CLI and log in first. See [Installing CLI](https://docs.omnistrate.com/getting-started/installing-ctl/index.md). - Review the generic [Account Onboarding](https://docs.omnistrate.com/getting-started/account-onboarding/index.md) page for the responsibility model and offboarding rules. - Review [BYOC Cloud Accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md) for operational guidance after onboarding. - For BYOA plans, make sure the target service plan is actually configured as BYOA before using `--customer-account-id` during deployment creation. ## Nebius-specific difference Nebius is the only provider in this guide set that uses **bindings** inside the account config. A binding is the combination of: - a Nebius project, - a Nebius service account, - a public key ID, and - the matching private key PEM. Each ready binding enables one Nebius region for deployment. Adding another region means adding another binding and replacing the full binding file during an update. # Specification Guide # Docker Compose Service Specification Omnistrate extends Docker Compose with `x-omnistrate-*` and related customer and internal integration tags. A Compose file remains the application topology, while these extensions describe the Plan, resources, customer inputs, infrastructure, lifecycle integrations, and control-plane behavior. ## What the Specification Covers A Compose service specification can define: - Services, images, ports, networks, volumes, configs, and secrets. - Plans, tenancy, deployment models, and account configuration. - API parameters, compute, storage, integrations, action hooks, and capabilities. - Variable interpolation and Omnistrate system parameters. Use the [Compose deployment strategy](https://docs.omnistrate.com/build-guides/compose-overview/index.md) to understand where Compose fits in the platform. For schema validation, native Compose tags, Omnistrate extensions, and complete configuration examples, continue to the [detailed Compose Service Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md). The detailed reference uses the same `x-` extension names used by the platform and is the source of truth for supported Compose configuration. # Choosing Your Service Specification Approach ## Service Specification Choice Overview This guide helps you understand when to use **Compose-based specifications** versus **Plan specifications** for defining your SaaS service in Omnistrate, with clear decision criteria and template references. For Hosted SaaS tenancy options, see the [Cellular Multi-Tenancy Overview](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-overview/index.md) and [Dedicated Tenancy Overview](https://docs.omnistrate.com/build-guides/dedicated-tenancy-overview/index.md). ## Quick Decision Matrix | Your Technology Stack | Recommended Approach | Template Reference | Specification Reference | | ------------------------ | --------------------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | | **Source code** | Compose-based Specification | [Compose Template](https://github.com/omnistrate-community/compose-based-template) | [Compose Specification](https://docs.omnistrate.com/spec-guides/compose-spec/index.md) | | **Container Image** | Compose-based Specification | [Compose Template](https://github.com/omnistrate-community/compose-based-template) | [Compose Specification](https://docs.omnistrate.com/spec-guides/compose-spec/index.md) | | **Docker Compose** | Compose-based Specification | [Compose Template](https://github.com/omnistrate-community/compose-based-template) | [Compose Specification](https://docs.omnistrate.com/spec-guides/compose-spec/index.md) | | **Helm Charts** | Plan Specification | [Plan Spec Template](https://github.com/omnistrate-community/plan-spec-template) | [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) | | **Kubernetes Operators** | Plan Specification | [Operator Spec Template](https://github.com/omnistrate-community/operator-spec-template) | [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) | | **Kustomize** | Plan Specification | [Plan Spec Template](https://github.com/omnistrate-community/plan-spec-template) | [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) | | **Terraform** | Plan Specification | [Plan Spec Template](https://github.com/omnistrate-community/plan-spec-template) | [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) | | **OpenTofu** | Plan Specification | [Plan Spec Template](https://github.com/omnistrate-community/plan-spec-template) | [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) | ## Compose Specification ### When to Use Compose Specification **✅ Choose Compose specification when:** - Your application is **containerized** and can be fully described using Docker Compose - You want to define your **entire service specification in a single file** - Your infrastructure requirements are **straightforward** (standard containers, volumes, networks) - You prefer **Docker Compose syntax** and workflows - You're building a **new containerized application** from scratch - You want **rapid prototyping** and simpler configuration management ### Key Characteristics of Compose Specification - **Single file approach**: Everything defined in `omnistrate-compose.yaml` - **Omnistrate extensions**: Uses `x-omnistrate-*` tags within the compose file - **Built-in parameter management**: API parameters defined directly in the compose spec - **Automatic infrastructure**: Omnistrate handles Kubernetes deployment automatically - **Simpler learning curve**: Familiar Docker Compose syntax with Omnistrate extensions ### Template and Examples for Compose Specification **📋 Template Repository:** [omnistrate-community/compose-based-template](https://github.com/omnistrate-community/compose-based-template) **Example structure:** ``` version: '3.9' x-omnistrate-service-plan: name: 'My Service Plan' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: web: image: nginx:latest ports: - "80:80" x-omnistrate-api-params: - key: instanceType description: Instance Type name: Instance Type type: String required: true x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: t3.medium ``` ### When Compose Specification is Best - **Web applications** with standard containerized architectures - **Microservices** that can be orchestrated with Docker Compose - **Development and testing** environments - **Simple to moderate complexity** deployments - **Teams familiar with Docker Compose** ______________________________________________________________________ ## Plan Specification ### When to Use Plan Specification **✅ Choose Plan specification when:** - You're using **Helm charts, Operators, Kustomize, or Terraform** - You need **complex Kubernetes configurations** beyond what Compose can express - You have **existing infrastructure as code** (Helm charts, Terraform modules) - You require **advanced Kubernetes features** (CRDs, operators, complex networking) - You need **multi-cloud Terraform deployments** - You want **separation of concerns** between service definition and plan configuration ### Key Characteristics of Plan Specification - **Separate specification file**: Plan defined in `spec.yaml`, separate from your infrastructure code - **Technology agnostic**: Works with Helm, Operators, Kustomize, and Terraform - **Advanced configurations**: Supports complex Kubernetes and cloud provider features - **Flexible deployment models**: Better support for BYOA (Bring Your Own Account) scenarios - **Enterprise features**: Advanced networking, security, and compliance capabilities ### Template and Examples for Plan Specification **📋 Template Repository:** [omnistrate-community/plan-spec-template](https://github.com/omnistrate-community/plan-spec-template) **Example structure:** ``` # spec.yaml name: "PostgreSQL Service" deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: "database" apiParameters: - key: "instanceType" description: "Database Instance Type" type: "string" required: true compute: instanceTypes: - name: "t3.medium" cloudProvider: "aws" helmChartConfiguration: chartName: "postgresql" chartVersion: "12.1.2" chartRepoName: "bitnami" chartRepoURL: "https://charts.bitnami.com/bitnami" chartValues: auth: postgresPassword: "$var.postgresPassword" ``` ### When Plan Specification is Best - **Helm-based applications** with complex chart configurations - **Kubernetes Operators** requiring custom resources and controllers - **Kustomize deployments** with sophisticated overlays and patches - **Terraform infrastructure** spanning multiple cloud providers - **Enterprise applications** with advanced security and compliance requirements - **Existing infrastructure as code** that you want to SaaS-ify ## Getting Started with Compose-based Specification 1. **Clone the template**: `git clone https://github.com/omnistrate-community/compose-based-template` 1. **Follow the guide**: [Build from Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) 1. **Reference documentation**: [Compose Specification](https://docs.omnistrate.com/spec-guides/compose-spec/index.md) ## Getting Started with with Plan Specification 1. **Clone the template**: `git clone https://github.com/omnistrate-community/plan-spec-template` 1. **Choose your technology**: 1. [Build from Helm](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) 1. [Build from Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md) 1. [Build from Kustomize](https://docs.omnistrate.com/getting-started/build-from-kustomize/index.md) 1. [Build from Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md) 1. **Reference documentation**: [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) For Kubernetes Operators, start from the dedicated operator template: ``` git clone https://github.com/omnistrate-community/operator-spec-template ``` ## Service Specification Choice Summary **Choose Compose-based specification** for straightforward containerized applications that fit well within Docker Compose paradigms. It offers simplicity and rapid development. **Choose Plan specification** for complex deployments using Helm, Operators, Kustomize, or Terraform, especially when you need advanced Kubernetes features or have existing infrastructure as code. Both approaches are fully supported by Omnistrate and can be used to build production-ready SaaS Product. The choice depends on your technology stack, complexity requirements, and team expertise. # Omnistrate Service Specification ## Omnistrate Service Specification Overview (Helm, Operators, Kustomize, Terraform, OpenTofu) When defining a service using Helm, Operators, Kustomize, Terraform or OpenTofu we need to define the details of the Plan using a specification file. Info When using docker compose there is no need to set up this file as the information for the Plan can be defined on the docker compose spec. More information under [Getting started / Setup using Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) A complete list of settings that can be configured on the Service Spec file is included below. The default path for a Plan specification file is `spec.yaml`. For Hosted SaaS tenancy options, see the [Cellular Multi-Tenancy Overview](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-overview/index.md) and [Dedicated Tenancy Overview](https://docs.omnistrate.com/build-guides/dedicated-tenancy-overview/index.md). ## Using Schema Validation To simplify the definition of this specification file Omnistrate provide a JSON schema that can be used for validation. You can use the JSON schema in IDEs that use the YAML Language Server (eg: VSCode / NeoVim). ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json ``` For [**IntelliJ**](https://www.jetbrains.com/help/idea/yaml.html#use-schema-keyword) replace the top line with the following line to set up the yaml schema ``` # $schema: https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json ``` ## Plan schema definitions This Service Schema defines the structure for a Plan specification in Omnistrate. It consists of several key components such as API parameters, capabilities, compute resources, network configurations, etc. Each component is detailed below. ### Root schema The root of the schema includes the following properties: | Property | Type | Description | | ----------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------- | | `name` | string | The name of the Plan. | | `deployment` | [deployment schema](#deployment-schema) | The deployment type of the Plan. The possible values are `hostedDeployment` and `byoaDeployment`. | | `features` | [features schema](#features-schema) | The list of features enabled for the Plan. The possible values are `INTERNAL` and `CUSTOMER`. | | `services` | array | A list of [service configurations](#service-schema) for the plan. **Required**. | | `pricing` | [pricing schema](#pricing-schema) | A list of pricing dimensions for the plan. | | `metering` | [metering schema](#metering-schema) | The configuration for metering. | | `billingProviders` | [billing provider schema](#billing-provider-schema) | A list of [billing providers](#billing-provider-schema). | | `maxNumberOfInstancesAllowed` | integer | The maximum number of instances allowed. | | `loadBalancers` | [load balancers schema](#load-balancers-schema) | The configuration for load balancers. | | `enableDeletionProtection` | boolean | Whether deletion protection is enabled for the plan. | ### Deployment schema The deployment type of the Plan. The possible values are `hostedDeployment` and `byoaDeployment`. Within the deployment type we need to add the information for the cloud provide account to use. If you are planning to use a single cloud provider, there is no need to add information for the others. | Property | Type | Description | | ---------------------------- | ------ | --------------------------------- | | `awsAccountId` | string | AWS Account ID. | | `awsBootstrapRoleAccountArn` | string | ARN of the AWS bootstrap role. | | `gcpProjectId` | string | GCP Project ID. | | `gcpProjectNumber` | string | GCP Project number. | | `gcpServiceAccountEmail` | string | Email of the GCP service account. | | `azureSubscriptionId` | string | Azure Subscription GUID. | | `azureTenantId` | string | Azure Tenant GUID. | | `ociTenancyId` | string | OCI Tenancy OCID. | | `ociDomainId` | string | OCI Domain OCID. | Note OCI is currently supported for **Custom Tenancy** deployments only (plan specification). Compose specification does not support OCI. ### Features schema Defines the available features for both internal and customer-facing use cases. Features are categorized into INTERNAL and CUSTOMER, specifying their respective functionalities. | Property | Type | Description | | ---------- | ------ | --------------------------------- | | `INTERNAL` | string | Features scoped to SaaS Providers | | `CUSTOMER` | string | Feature scoped to end customer | #### Supported INTERNAL Features | Property | Type | Description | | -------- | ------ | ---------------------------------------------------------- | | `logs` | object | Enable SaaS Provider logs on the operational log dashboard | #### INTERNAL.logs schema Use `features.INTERNAL.logs` to enable SaaS Provider logs. An empty object enables Omnistrate-managed logs: ``` features: INTERNAL: logs: {} ``` To send Helm resource internal logs to CloudWatch Logs, set `provider: native`: ``` features: INTERNAL: logs: provider: native ``` | Property | Type | Description | | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------- | | `provider` | string | Optional. Set to `native` to send Helm resource internal logs to CloudWatch Logs. Omit this field to use Omnistrate-managed logs. | For provider-hosted Helm deployments, native logging is currently supported only when the instance is deployed on AWS. Logs are sent to CloudWatch Logs in the Hosted SaaS Provider's AWS deployment account using the deployment cell's native AWS credentials. For Helm resources in BYOC deployments, including BYOC On-Premise, internal native logs are sent to CloudWatch Logs in the Hosted SaaS Provider's AWS account regardless of the dataplane cloud. Omnistrate provides scoped, temporary AWS credentials that allow writes to the instance's CloudWatch log groups and stream prefix. For this Helm configuration, the BYOC dataplane must have an HTTPS network path to the regional AWS CloudWatch Logs endpoint. It can use public egress or private, routed connectivity to an AWS CloudWatch Logs interface VPC endpoint. Cross-cloud and on-premise private paths require appropriate routing and DNS, such as connectivity through a VPN or AWS Direct Connect. Native logs are not supported when the dataplane cannot reach CloudWatch Logs. This setting applies to `INTERNAL` logs for the SaaS Provider operations view. For Helm configuration and account requirements, see [Send Helm Resource Logs to CloudWatch Logs](https://docs.omnistrate.com/getting-started/build-from-helm/#send-helm-resource-logs-to-cloudwatch-logs). #### Supported CUSTOMER Features | Property | Type | Description | | ----------- | ------ | -------------------------------------------------------------------------------------- | | `logs` | string | Enable customer visible logs | | `metrics` | object | Enable customer-visible metrics collection, additional metric modeling, and dashboards | | `licensing` | string | Enable software license management | #### CUSTOMER.metrics schema (advanced metrics and dashboards) Use `features.CUSTOMER.metrics` to define how customer-facing metrics are collected and visualized. ``` features: CUSTOMER: metrics: scrapeTargets: - name: postgres podSelector: app.kubernetes.io/name: postgresql endpoint: portNumber: 9187 path: /metrics scheme: http additionalMetrics: postgres: metrics: pg_stat_activity_count: "Active Sessions": aggregationFunction: sum labelFilters: datname: postgres state: active pg_locks_count: "Locks by Mode": aggregationFunction: sum groupByLabels: - mode labelFilters: datname: postgres dashboards: exporter: title: "Postgres Exporter" sections: - title: "Sessions and Locks" panels: - title: "Active Sessions" type: timeseries unit: short targets: - ref: "Active Sessions" - title: "Locks by Mode" type: timeseries unit: short targets: - ref: "Locks by Mode" ``` Configuration structure and intent: - `scrapeTargets`: defines where metrics are scraped from. - Use `podSelector` and/or `serviceSelector` to discover endpoints. - `endpoint` supports `portName` or `portNumber`, plus `path`, `scheme`, and optional query `params`. - `additionalMetrics..metrics`: models raw metrics into stable named series. - `` should match your service component name (for example, `postgres`). - First key is raw metric name (for example, `pg_stat_activity_count`). - Nested key is exported series name (for example, `"Active Sessions"`), referenced by dashboards. - Optional per-series controls: - `aggregationFunction`: `sum`, `min`, `max`, `avg` - `groupByLabels`: preserves selected dimensions in the output - `labelFilters`: filters label key/value pairs before aggregation - `additionalMetrics..dashboards`: builds custom dashboards from modeled series. - `title`: dashboard title override - `sections[]`: ordered dashboard sections - `panels[]`: panel definitions (`timeseries`, `stat`, `gauge`, `table`) - `targets[]`: - `ref`: reference a modeled series name from `metrics` - `expr`: raw PromQL (optional override) - `legend`, `axis`: optional display controls - Optional panel fields: `description`, `unit`, `min`, `max`, `thresholds`, `tags` Dashboard property details: | Path | Type | Required | Description | | ------------------------------------ | ------------- | -------------------- | -------------------------------------------------------------- | | `dashboards..title` | string | No | Display title for the dashboard | | `dashboards..sections` | array | Yes | Ordered list of sections | | `sections[].title` | string | Yes | Section heading | | `sections[].description` | string | No | Optional section help/context | | `sections[].panels` | array | Yes | Panels under section | | `panels[].title` | string | Yes | Panel title | | `panels[].type` | enum | Yes | `timeseries`, `stat`, `gauge`, `table` | | `panels[].description` | string | No | Optional panel description | | `panels[].unit` | string | No | Unit formatting (for example `bytes`, `s`, `short`) | | `panels[].min` / `panels[].max` | number | No | Explicit panel bounds | | `panels[].targets` | array | Yes | Query targets for panel | | `targets[].ref` | string | Recommended | Modeled series key from `additionalMetrics..metrics` | | `targets[].expr` | string | Optional alternative | Raw PromQL expression | | `targets[].legend` | string | No | Legend override | | `targets[].axis` | enum | No | Axis side: `left` or `right` | | `panels[].thresholds` | array | No | Threshold definitions | | `thresholds[].color` | string | Yes (if used) | Threshold color | | `thresholds[].value` | number | Yes (if used) | Threshold boundary value | | `panels[].tags` | array[string] | No | Optional panel tags | Practical guidance: - Start by validating `scrapeTargets` first; missing scrape configuration is the most common reason custom panels stay empty. - Use modeled series names in `targets.ref` so dashboard layout stays stable if raw metric labels change. - Apply `labelFilters` before aggregation to avoid mixing unrelated dimensions. - For copy/paste panel templates (health stat, saturation trend, distribution table), see [Build Guide / Integrations](https://docs.omnistrate.com/build-guides/integrations/#custom-metrics-and-dashboards-advanced). - For legacy-to-advanced migration steps, see [Build Guide / Integrations](https://docs.omnistrate.com/build-guides/integrations/#migration-guide-legacy-custom-metrics-to-advanced-metrics-dashboards). ### Service schema Each Plan has a list of Resources that need to be created as part of the Plan. For each service component we need to define a set of properties. | **Property** | **Type** | **Description** | | -------------------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `name` | string | The name of the Resources. | | `internal` | boolean | Defines if the Resource can be created by customers or is an internal resource used by other Resources | | `dependsOn` | array | List of Resources defined on this service that need to be created before this Resource | | `apiParameters` | [api parameters schema](#api-parameters-schema) | List of parameters for the Resource. | | `compute` | [compute schema](#compute-schema) | The compute resources allocated for the service. | | `network` | [network schema](#network-schema) | Network-related configuration for the service, including ports. | | `capabilities` | [capabilities schema](#capabilities-schema) | Defines Resource characteristics like running on more than one AZ. | | `helmChartConfiguration` | [helm schema](#helm-chart-configuration-schema) | Helm chart configuration, if the Resource is defined using Helm. Can be specified for each Helm resource in the service | | `operatorCRDConfiguration` | [operator schema](#operator-crd-schema) | Operator CRD configuration for provisioning the Resource, if the service is component is defined as an Operator. Can be specified for each Operator resource in the service | | `kustomizeConfiguration` | [kustomize schema](#kustomize-configuration-schema) | Kustomize configuration for provisioning the Resource if the Resource is defined using Kustomize. Can be specified for each Kustomize resource in the service | | `terraformConfigurations` | [terraform schema](#terraform-configuration-schema) | Terraform configuration for provisioning the Resource on each cloud provider, if the Resource requires a terraform stack. Can be specified for each Terraform/OpenTofu resource in the service | | `deploymentTarget` | [deployment target schema](#deployment-target-schema) | Controls where the Resource executes. Use this for resources that should run in the control plane account instead of the deployment cell account. | ### `API parameters schema` Defines the list of API parameters for the Resource | Property | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `key` | string | The key of the parameter. | | `description` | string | Description of the parameter. | | `name` | string | The display name of the parameter. | | `type` | string | The type of the parameter (e.g., string). | | `modifiable` | boolean | Indicates if the parameter is modifiable. | | `required` | boolean | Indicates if the parameter is required. | | `export` | boolean | If true, the parameter is exportable. | | `defaultValue` | string | Default value of the parameter. | | `regex` | string | A regular expression to validate the parameter value. | | `options` | array | Available options for the parameter. | | `labeledOptions` | object | A set of labeled options as key-value pairs. | | `limits` | object | Limit settings for the parameter. | | `dependentResourceKey` | string | Key for dependent resources. | | `parameterDependencyMap` | object | Dependency map for other parameters. | | `scope` | object | Conditions to restrict parameter visibility by cloud provider. Contains `CloudProviders` (array of provider keys such as `aws`, `gcp`, `azure`, or `nebius`). | The valid types are described on the [api parameters build guide](https://docs.omnistrate.com/build-guides/api-params/#parameter-types) See [Scoped API Parameters](https://docs.omnistrate.com/build-guides/api-params/#scoped-api-parameters) for usage examples. ### `Compute schema` Configuration for compute resources for the Resource. | Property | Type | Description | | ------------------ | ------- | -------------------------------------- | | `instanceTypes` | array | List of instance types supported. | | `rootVolumeSizeGi` | integer | The size of the root volume in Gi. | | `cpuArchitecture` | string | The CPU architecture for the resource. | #### Instance Type schema | Property | Type | Description | | ------------------------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `name` | string | The name of the instance type. | | `apiParam` | string | The API parameter name associated with the instance type. | | `cloudProvider` | string | The cloud provider for the instance type (e.g., AWS, GCP, Azure, OCI, Nebius). | | `platform` | string | Nebius-only compute platform for this instance type. Required when `cloudProvider` is `nebius`. | | `configurationOverrides` | object | Optional per-instance-type overrides such as OS family, root volume overrides, accelerator settings, labels, taints, or Nebius GPU cluster placement. See [configurationOverrides](#configurationoverrides). | #### configurationOverrides The `configurationOverrides` object allows you to customize node-level settings for specific instance types. Available overrides: | Property | Type | Description | | -------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `OsFamily` | string | The operating system family for nodes. Set to `amazonlinux` for Amazon Linux 2023 on AWS. | | `acceleratorConfiguration` | object | GPU accelerator attachment for GCP N1 instances with Tesla GPUs. See the [GPU Accelerator Configuration Guide](https://docs.omnistrate.com/infra-guides/gpu-accelerator-configuration/index.md) for details. | | `GpuClusterID` | string | Nebius-only. The GPU cluster ID for placing the instance in a specific Nebius GPU cluster. | ##### OsFamily Set `OsFamily` to override the default operating system used for Kubernetes nodes. Currently supported values: | Value | OS | Cloud provider | | ------------- | ----------------- | -------------- | | `amazonlinux` | Amazon Linux 2023 | AWS | Nebius-specific notes: - For `cloudProvider: nebius`, the selected `name` or `apiParam` resolves to a Nebius preset. - `platform` is required for Nebius instance-type entries. - `configurationOverrides.GpuClusterID` is optional and Nebius-only. Use it when the preset must land in a specific Nebius GPU cluster. **Example — Nebius GPU cluster:** ``` compute: instanceTypes: - name: 8gpu-128vcpu-1600gb cloudProvider: nebius platform: gpu-h100-sxm configurationOverrides: GpuClusterID: computegpucluster-e00abcxfadyhcvpkdx ``` **Example — Amazon Linux 2023:** ``` services: - name: my-service compute: instanceTypes: - name: t3.medium cloudProvider: aws configurationOverrides: OsFamily: amazonlinux ``` ##### acceleratorConfiguration Attach external Tesla GPUs to GCP N1 instances. Required properties: | Property | Type | Description | | -------- | ------- | ------------------------------------------------------------------------ | | `type` | string | The GPU accelerator type (e.g., `nvidia-tesla-t4`, `nvidia-tesla-v100`). | | `count` | integer | Number of GPUs to attach (minimum: 1). | **Example — Tesla T4 on N1:** ``` services: - name: gpu-service compute: instanceTypes: - name: n1-standard-8 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-t4" count: 1 ``` For the full list of supported GPU types and built-in GPU instances that do not require `acceleratorConfiguration`, see the [GPU Accelerator Configuration Guide](https://docs.omnistrate.com/infra-guides/gpu-accelerator-configuration/index.md). Note Changing `configurationOverrides` (such as switching the OS family) creates new node pools. The old node pools are not deleted immediately. Review node pool quota before upgrading, and clean up stale node pools afterward. See [Node Pool Cleanup](https://docs.omnistrate.com/operate-guides/deployment-cell-nodepools/#node-pool-cleanup). ### Network schema Configuration for network interfaces for the Resource. | Property | Type | Description | | ------------ | ----- | -------------------------------------------------------------------------------------------------------------------------------------------- | | `ports` | array | A list of ports and port ranges. Supports integer values (e.g., `5432`) and range syntax (e.g., `8000-8100`). | | `namedPorts` | map | A map of port names to ports or port ranges. Supports integer values and range syntax (e.g., `postgresql: 5432`, `grpc-range: 50000-50010`). | ### Deployment target schema Controls where a Resource executes. | Property | Type | Description | | --------- | ------ | ------------------------------------------------------------------------------------------------------------------------------ | | `account` | string | Target account for the Resource. Use `ControlPlane` for resources that should execute in the Omnistrate control plane account. | **Example:** ``` services: - name: controlPlaneInfra internal: true deploymentTarget: account: ControlPlane terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/control_plane gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` ### Capabilities schema Specifies capabilities enabled for the Resource. | Property | Type | Description | | ----------------- | ------- | -------------------------------------- | | `enableMultiZone` | boolean | Whether multi-zone support is enabled. | ### Helm chart configuration schema When defining a Resource using Helm charts we need to specify the Helm chart definition and values Info You can use system parameters to customize Helm Chart values. A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). | Property | Type | Description | | ---------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `chartName` | string | Name of the Helm chart. Required for repository-backed charts. For local artifacts, Omnistrate reads this from `Chart.yaml`. | | `chartVersion` | string | Version of the Helm chart. Required for repository-backed charts. For local artifacts, Omnistrate reads this from `Chart.yaml`. | | `chartRepoName` | string | Name of the chart repository. Required for repository-backed charts. | | `chartRepoURL` | string | URL of the chart repository. Required for repository-backed charts. | | `artifactRelativePath` | string | Relative path to a local chart directory or packaged `.tar.gz` / `.tgz` chart. Use with `omnistrate-ctl build`; do not combine with `chartRepoName` / `chartRepoURL`. | | `chartValues` | object | Values to be passed to the Helm chart. | | `chartAffinityControl` | object | Affinity control settings for the Helm chart. See [chart affinity control schema](#chart-affinity-control-schema). | | `authProvider` | object | Username and password to access a private Helm chart repository. Omit this field for private Amazon ECR OCI chart repositories that should use AWS IAM credential resolution. | For private Amazon ECR OCI chart repositories, set `chartRepoURL` to the ECR OCI URL and omit `authProvider`. Omnistrate resolves short-lived ECR credentials at deployment time when AWS account setup allows ECR Helm chart pulls. ### Chart affinity control schema Controls automatic injection of Kubernetes affinity rules into Helm chart manifests. By default, Omnistrate automatically injects node affinity and pod anti-affinity rules to ensure proper scheduling on Omnistrate-managed nodes. For more details, see [Helm Chart Customization](https://docs.omnistrate.com/build-guides/helm-charts-customize/#automatic-affinity-injection-default). | **Property** | **Type** | **Default** | **Description** | | ------------------ | -------- | ----------- | --------------------------------------------------------------------------------------------------- | | `enableInjection` | boolean | `true` | Enables automatic injection of node affinity and pod anti-affinity rules into Helm chart manifests. | | `enableSharedHost` | boolean | `true` | Enables shared host support, allowing pods to be co-located when appropriate. | ### Operator CRD schema When defining a Resource using Operators CRD we need to specify the CDR information as test and add a reference to the required Helm charts Info You can use system parameters to customize Operators CRD. A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). Tip For operator-backed services with multi-step lifecycle operations, prefer `systemWorkflows` on the root service resource. This lets you model create, modify, delete, start, stop, capacity changes, backup, restore, and delete-backup with Argo Workflow-style DAGs. Use `customWorkflows` only for provider-defined operations outside the standard platform lifecycle APIs. See [Build with Kubernetes Operators](https://docs.omnistrate.com/build-guides/operators/index.md) and the [operator spec template](https://github.com/omnistrate-community/operator-spec-template). Note `operatorCRDConfiguration.template`, `operatorCRDConfiguration.supplementalFiles`, and `operatorCRDConfiguration.readinessConditions` are deprecated for operator lifecycle management. Use `systemWorkflows` for lifecycle resources, readiness, failure conditions, and outputs. `operatorCRDConfiguration.helmChartDependencies` remains supported for installing the operator and CRDs. | **Property** | **Type** | **Description** | | ----------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------- | | `template` | string | Deprecated. The template for the operator CRD configuration. Prefer `systemWorkflows`. | | `readinessConditions` | object | Deprecated. Conditions that determine the readiness state of the CRD. Prefer workflow `successCondition` and `failureCondition`. | | `supplementalFiles` | array | Deprecated. A list of supplemental files for the CRD. Prefer workflow-managed resources. | | `outputParameters` | object | A map of output parameters with string values. | | `helmChartDependencies` | array | A list of [Helm chart dependencies](#helm-chart-dependency-schema) referenced by the CRD. | #### Helm chart dependency schema | **Property** | **Type** | **Description** | | --------------- | -------- | ------------------------------------------------------- | | `chartName` | string | The name of the Helm chart dependency. | | `chartVersion` | string | The version of the Helm chart dependency. | | `chartRepoName` | string | Name of the chart repository. | | `chartRepoURL` | string | URL of the chart repository. | | `chartValues` | object | Values to be passed to the Helm chart. | | `authProvider` | object | Username and password to access private Helm cart repo. | ### Kustomize configuration schema When defining a Resource using Kustomize you need to specify a reference to the GitHub repository where the Kustomize stack is defined. | **Property** | **Type** | **Description** | | ------------------ | ---------------------------------------------- | -------------------------------------------------------- | | `kustomizePath` | string | The path to the Kustomize configuration. | | `gitConfiguration` | [git access schema](#git-configuration-schema) | Configuration for Git access related to Kustomize files. | ### Terraform configuration schema When defining a Terraform (or OpenTofu) stack to be executed to provision the Resource, you need to add a reference to the GitHub repository where the stack definition is stored. A different Terraform stack can be defined for each cloud provider using the `configurationPerCloudProvider` property. Info You can use system parameters to pass input parameters to your Terraform stack. A detailed list of system parameters can be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). | **Property** | **Type** | **Description** | | ------------------------------- | ------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | | `configurationPerCloudProvider` | [per-cloud-provider schema](#terraform-per-cloud-provider-configuration-schema) | Terraform configuration for each cloud provider. Accepts `aws`, `gcp`, `azure`, `oci`, and `nebius` keys. | ### Terraform per-cloud-provider configuration schema Each key under `configurationPerCloudProvider` (`aws`, `gcp`, `azure`, `oci`, `nebius`) accepts the following properties: | **Property** | **Type** | **Description** | | ----------------------------- | --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `terraformPath` | string | The directory path within the repository containing the Terraform stack for this cloud provider. **Required**. | | `serviceAccountID` | string | Nebius-only Terraform service account ID. Required when the cloud provider key is `nebius`. | | `publicKeyID` | string | Nebius-only Terraform public key ID. Required when the cloud provider key is `nebius`. | | `privateKeyPEM` | string | Nebius-only Terraform private key PEM. Required when the cloud provider key is `nebius`. Supports `$secret.` references. | | `requiredOutputKeys` | array[string] | List of Terraform output keys that must be present after apply. | | `requiredOutputs` | [terraform output schema](#terraform-output-schema) | List of Terraform outputs to validate and optionally export from the Resource. | | `variablesValuesFileOverride` | string | Inline Terraform variable definitions in `.tfvars` format. Supports Omnistrate system parameters for dynamic value injection at deployment time. | | `cliConfigFileOverride` | string | Inline [OpenTofu CLI configuration](https://opentofu.org/docs/cli/config/config-file/) content applied to every Terraform operation for this cloud provider via `TF_CLI_CONFIG_FILE` — for example, provider installation mirrors or credentials for private registries. Supports Omnistrate system parameters. Redacted in API responses for read-only roles. | | `gitConfiguration` | [git access schema](#git-configuration-schema) | Configuration for Git access to the repository containing the Terraform stack. Required when `artifactRelativePath` is not set. | | `artifactRelativePath` | string | Relative path to a local Terraform module directory. Use with `omnistrate-ctl build`; do not combine with `gitConfiguration`. | Provider-specific notes: - For Hosted SaaS deployments, Omnistrate manages Terraform execution identities. On AWS, GCP, Azure, and OCI, Omnistrate auto-creates the identity. Assign the permissions your Terraform stack requires to the identity. You do not need to specify the identity in the Plan spec. - `configurationPerCloudProvider.nebius` requires `serviceAccountID`, `publicKeyID`, and `privateKeyPEM`. - Prefer `$secret.` for `privateKeyPEM` so the PEM stays out of source control. **Example:** ``` terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws requiredOutputs: - key: database_endpoint exported: true variablesValuesFileOverride: | vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git gcp: terraformPath: /terraform/gcp gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git nebius: terraformPath: /terraform/nebius serviceAccountID: serviceaccount-e00vqdp9fskhmmaan8 publicKeyID: publickey-e00h9scsyy9mbefrjf privateKeyPEM: $secret.nebiusTerraformPrivateKey gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` For more details, see [Terraform Overview](https://docs.omnistrate.com/build-guides/terraform-overview/index.md), [Multi-Cloud Configuration](https://docs.omnistrate.com/build-guides/terraform-multi-cloud/index.md), and [Input Parameters and Output Mapping](https://docs.omnistrate.com/build-guides/terraform-params-outputs/index.md). ### Terraform output schema Defines a Terraform output that Omnistrate should validate after apply and optionally export from the Resource. | Property | Type | Description | | ---------- | ------- | -------------------------------------------------------------------------------------------- | | `key` | string | Terraform output key to validate. | | `exported` | boolean | If `true`, surface this output as an exported field on the Terraform Resource in Omnistrate. | ### Git configuration schema Defines the configuration for Git access. | Property | Type | Description | | --------------- | ------ | -------------------------------------------------------------- | | `userName` | string | Git user name. | | `accessToken` | string | Access token for Git authentication. | | `reference` | string | Reference to the Git branch, tag, or commit SHA. **Required**. | | `repositoryUrl` | string | URL of the Git repository. **Required**. | The `reference` field accepts branch names, tags, and full commit SHAs. Using a specific commit SHA pins your deployment to an exact version of the code, which is useful for reproducible builds and controlled rollouts. ``` git: # Using a branch reference: refs/heads/main repositoryUrl: https://github.com/my-org/my-repo ``` Using a tag: ``` reference: refs/tags/v1.2.0 repositoryUrl: https://github.com/my-org/my-repo ``` Using a specific commit SHA: ``` reference: a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2 repositoryUrl: https://github.com/my-org/my-repo ``` Tip Pinning to a commit SHA ensures that your deployment uses the exact same code every time, even if a branch or tag is updated. This is recommended for production deployments where reproducibility is critical. ### Pricing schema Defines the pricing structure for the plan based on various dimensions. | Property | Type | Description | | ----------- | ------ | ----------------------------------------------------- | | `dimension` | string | The name of the pricing dimension (e.g., "CPU"). | | `unit` | string | (Optional) The unit of measurement for the dimension. | | `timeUnit` | string | (Optional) The time unit for billing (e.g., "hour"). | | `price` | number | The price per unit. | Info For more detailed information on the **pricing**, **metering**, **billingProviders** configuration, please see [End-to-End Billing](https://docs.omnistrate.com/fin-ops-guides/billing/index.md) and [Usage Metering](https://docs.omnistrate.com/fin-ops-guides/metering/index.md). ### Metering schema Defines the configuration for metering. | Property | Type | Description | | ---------------- | ------ | -------------------------------------- | | `s3BucketArn` | string | The ARN of the S3 bucket for metering. | | `s3BucketRegion` | string | The region of the S3 bucket. | | `gcsBucketName` | string | The name of the GCS bucket. | ### Billing Provider schema Defines a billing provider. | Property | Type | Description | | -------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `name` | string | The name of the billing provider. Use `stripe` (case-insensitive) for Stripe integration or the name you set for your bring-your-own (BYO) billing provider. | | `externalProductID` | string | (Stripe only) The external product ID. | | `enablePaywall` | boolean | (Stripe only) Specifies if a paywall is enabled. | | `disableInvoiceGeneration` | boolean | (Stripe only) Specifies if Stripe invoice generation is disabled for this plan. | | `isDefault` | boolean | Specifies if this is the default provider. | ### Load Balancers schema Defines the configuration for L7 and L4 load balancers. | Property | Type | Description | | -------- | ----- | ------------------------------------------------------------------------------------ | | `https` | array | A list of [L7 load balancer configurations](#l7-load-balancer-configuration-schema). | | `tcp` | array | A list of [L4 load balancer configurations](#l4-load-balancer-configuration-schema). | #### L7 Load Balancer Configuration schema | Property | Type | Description | | ----------------- | ------- | -------------------------------------------------------------------------------- | | `name` | string | The name of the L7 load balancer. | | `description` | string | A description of the L7 load balancer. | | `enableCustomDNS` | boolean | Specifies if custom DNS is enabled. | | `paths` | array | A list of [L7 load balancer paths](#l7-load-balancer-path-configuration-schema). | #### L7 Load Balancer Path Configuration schema | Property | Type | Description | | ----------------------------- | ------- | ------------------------------------------ | | `targetKubernetesServiceName` | string | The name of the target Kubernetes service. | | `associatedResourceKey` | string | The associated resource key. | | `path` | string | The path for the load balancer rule. | | `backendPort` | integer | The backend port. | #### L4 Load Balancer Configuration schema | Property | Type | Description | | ------------- | ------ | -------------------------------------------------------------------------------- | | `name` | string | The name of the L4 load balancer. | | `description` | string | A description of the L4 load balancer. | | `ports` | array | A list of [L4 load balancer ports](#l4-load-balancer-port-configuration-schema). | #### L4 Load Balancer Port Configuration schema | Property | Type | Description | | ------------------------ | ------- | ----------------------------------- | | `associatedResourceKeys` | array | A list of associated resource keys. | | `backendPort` | integer | The backend port. | | `ingressPort` | integer | The ingress port. | # Build Guides # Action hooks ## What is an Action Hook Action hooks allows you to customize your control plane by injecting custom code at different phase of the lifecycle of your SaaS operations ``` x-omnistrate-actionhooks: - scope: CLUSTER type: INIT commandTemplate: > PGPASSWORD={{ $var.postgresqlRootPassword }} psql -U postgres -h writer {{ $var.postgresqlDatabase }} -c "create extension vector" ``` As an example, in the above case, we are enabling vector extension for Postgres on cluster initialization. ## Example use cases for Action Hooks Action hooks can be used for wide variety of use-cases. Here are some examples: - To enable extensions on cluster creation - To detect process deadlatches - To run rebalance command after adding or removing a node - To run post upgrade action ## Action hooks types Action hooks are categorized into two scopes: - Node: for every node of type Resource, we will run the associated action hook - Cluster: for every Resource instance, we will run the associated action hook ### Node scope - **`HEALTH_CHECK`**: Runs periodic health checks on the node to ensure it's functioning correctly. This is mapped to Kubernetes liveness probes and is useful for detecting when a service becomes unresponsive and needs a restart.\ By default, a TCP ping is performed on the service port (if present). You can override this behavior by defining a custom action hook.\ *(Optional: Define only if liveness probing is required.)* - **`READINESS_CHECK`**: Determines whether the service is ready to accept traffic. This is mapped to Kubernetes readiness probes. If not explicitly defined, it defaults to `HEALTH_CHECK`.\ *(Optional: Use when readiness conditions differ from general health status.)* - **`STARTUP_CHECK`**: Ensures the service has fully started before it begins accepting traffic. This is mapped to Kubernetes startup probes. If not explicitly defined, it defaults to `HEALTH_CHECK`.\ *(Optional: Useful for applications with long initialization time, for example a large backup restore operation)* - **`POST_START`**: Executes a post-start action on the node after it starts.\ *(Optional: Can be used for setup tasks like registering with an external system.)* - **`ADD`**: Triggers an action when a new node is added.\ *(Optional: Useful for tasks like joining a cluster or initializing configurations.)* - **`REMOVE`**: Triggers an action when a node is removed.\ *(Optional: Useful for cleanup tasks like de-registering from a service registry.)* - **`INIT`**: Runs an initialization action before a node starts.\ *(Optional: Can be used for pre-start setup like loading configurations or initializing dependencies.)* - **`PROMOTE`**: Triggers an action when a node is promoted recovered and can be promote to a primary role.\ *(Optional: Useful for tasks like promoting a replica to a leader role.)* - **`DEMOTE`**: Triggers an action when a node is deemed unhealthy, before replacing the node.\ *(Optional: Useful for tasks like demoting a leader role to a replica and relinquish leadership.)* - **`PRE_UPGRADE`**: Runs before the node is restarted during upgrades or machine type changes.\ *(Optional: Useful for draining connections, flushing data to disk, or gracefully removing the node from a cluster before it restarts.)* - **`POST_UPGRADE`**: Runs after the node is restarted during upgrades or machine type changes.\ *(Optional: Useful for re-joining a cluster, validating the node is healthy after restart, or restoring connections.)* ### Cluster scope - **`POST_START`**: Executes a post-start action after all nodes have started for a given Resource.\ *(Optional: Can be used for tasks like finalizing configurations, notifying external systems, or running post-deployment scripts.)* - **`PRE_START`**: Runs a pre-start action before all nodes are started for a given Resource.\ *(Optional: Useful for preparing shared resources, validating configurations, or performing pre-initialization tasks.)* - **`POST_STOP`**: Executes a post-stop action after all nodes are stopped for a given Resource.\ *(Optional: Can be used for cleanup tasks, releasing resources, or notifying dependent services.)* - **`PRE_STOP`**: Runs a pre-stop action before all nodes are stopped for a given Resource.\ *(Optional: Useful for gracefully shutting down services, draining connections, or backing up critical data.)* - **`POST_UPGRADE`**: Executes a post-upgrade action after all nodes are upgraded for a given Resource. This hook also fires after machine type changes.\ *(Optional: Can be used for verifying upgrade success, running data migrations, or reloading configurations.)* - **`PRE_UPGRADE`**: Runs a pre-upgrade action before all nodes are upgraded for a given Resource. This hook also fires before nodes are restarted for a machine type change.\ *(Optional: Useful for backing up data, validating system state, or ensuring compatibility before upgrading.)* - **`INIT`**: Runs an initialization action before all nodes are started for a given Resource.\ *(Optional: Can be used for configuring dependencies, or preparing the system for operation.)* If you have a requirement to add action hooks at other stages/operations, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) ## Action hooks runtime environment Action hooks are run as independent pods to minimize any runtime impact on the other data plane (aka application) resources. Having said that, action hooks can interact with other resources to perform its function. Action hooks are defined based on when you want to run them, i.e. defining an action on Resource doesn’t mean that it will be executed inside that component pod. Instead, it will be run based on the lifecycle of the Resource as defined by the action type. In the above example, action hook will be run when cluster resource is first initialized. Let's say tenant #2 is getting initialized and as you can see action hook is getting run as an independent pod alongside Resource pod: ## Action hooks programming model We support Bash scripts as a runtime environment. You can also invoke any third-party service or function directly from action hooks. If you would like to see a Python-based runtime environment or have other suggestions, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) with more details about your use case. ## Action hooks variables You can inject any system or dynamic variables into your action hook. For the full list, please see [here](https://docs.omnistrate.com/build-guides/api-params/#system-parameters) and [here](https://docs.omnistrate.com/build-guides/api-params/#dynamic-parameters). ## Action hooks lifecycle In this section, we will discuss how different action hooks types are run for different SaaS operations. To illustrate it, let's say we are building a service with Cluster resource containing two Resources: `SC1` and `SC2`. Here is how a dependency structure may look like: Now, if you have to define different action hook types on the Cluster resource, here is how they will get executed for each of the SaaS operations: - Provisioning operation: As an example, you can see `init` action hook is run after provisioning of all the Resources associated with the Cluster resource. - Start/Restart operation - Upgrade operation - Stop operation - Add capacity operation - Remove capacity operation - Deprovisioning operation ## How to configure action hooks? Use `x-omnistrate-actionhooks` compose tag. For more details on compose tags, see [here](https://docs.omnistrate.com/build-guides/compose-spec/#custom-tags) Alternatively, you can use register action hook API [here](https://api.omnistrate.cloud/docs/external/#tag/resource-api/operation/resource-api#RegisterActionHook) or follow our intuitive UI to configure your hooks. ## Action hooks examples ### Kernel Parameter Tuning example In this example, we are using a NODE/INIT action hook to tune kernel parameters. This is restricted to Plans with dedicated tenancy type where the workload is isolated to its own VM: ``` services: high-performance-db: image: postgres:15 x-omnistrate-actionhooks: - image: busybox:1.37 scope: NODE type: INIT command: - /bin/sh - -c commandTemplate: | # Set the AIO max number of events sysctl -w fs.aio-max-nr='102052' ``` ### Scaling example In this example, we are using action hooks to add a node to MongoDB cluster on scale up: ``` services: mongodb-primary: image: docker.io/omnistrate/mongodb:6.0-3 x-omnistrate-actionhooks: - scope: NODE type: ADD commandTemplate: | #!/bin/bash set -ex # Check if NODE_NAME is not equal to 'mongodb-primary-0' if [ "$NODE_NAME" != {{ $sys.compute.nodes[0].name }} ]; then # Run the mongosh command mongosh "mongodb://{{ $var.mongodbUsername }}:{{ $var.mongodbPassword }}@{{ $sys.compute.nodes[0].name }}:27017/?authMechanism=DEFAULT" --eval "rs.add( { host: '{{ $sys.compute.node.name }}' } )" fi ``` ### Health check example In this example, we are using action hooks to configure deep process liveness check on Couchbase Server. ``` #!/bin/sh IP=${IP:=127.0.0.1} PORT=${PORT:=8091} QUERY_PORT=${QUERY_PORT:=8093} USERNAME=${USERNAME:="Administrator"} PASSWORD=${PASSWORD:="password"} N1QL_STMT='SELECT name FROM `travel-sample`.inventory.hotel LIMIT 1' i=1 numbered_echo() { echo "[$i] $@" i=$(expr $i + 1) } # Check if couchbase server is up check_db() { curl --silent --output /dev/null http://$IP:$PORT echo $? } SUCCESS_CODE=200 EXIT_CODE=0 response="" # Issue POST and extract response launch_query() { # final_cmd="curl --silent $QUERY" numbered_echo "curl --write-out %{http_code} --output /dev/null --silent -u $USERNAME:$PASSWORD http://$IP:$QUERY_PORT/query/service -d \"statement=$N1QL_STMT\"" response=$(curl --write-out %{http_code} --output /dev/null --silent -u $USERNAME:$PASSWORD http://$IP:$QUERY_PORT/query/service -d "statement=$N1QL_STMT") echo $response if [ $response -ne $SUCCESS_CODE ] then EXIT_CODE=1 echo "Received: $response Expected: $SUCCESS_CODE => Exiting with exit code: $EXIT_CODE" exit $EXIT_CODE fi } # Wait until cb db is ready until [[ $(check_db) = 0 ]]; do numbered_echo "cb db not available yet" sleep 1 done numbered_echo "cb db is available" # LAUNCH QUERY launch_query exit $EXIT_CODE ``` # Air-Gapped Deployment ## What Is an Air-Gapped Deployment An air-gapped deployment Plan allows you to package your Helm chart as a self-contained installer that your customers can deploy onto their own Kubernetes clusters. The Plan specification defines deployment requirements, API parameters, lifecycle hooks, and Helm chart configuration. Note An installer Plan can include multiple Helm chart resources and multiple container image registry copy resources. Model each Helm release as its own `services[]` entry, and model each source registry or repository copy path as its own internal image sync service. Note The installer assumes that the customer running it has `kubectl`, `helm`, and a valid `kubeconfig` for the target cluster. ## Configuring the deployment block Set the product name and the air-gapped deployment requirements: ``` name: My Application deployment: requirements: k8sVersion: ">=1.30.0" # optional onPremDeployment: # AWS account that hosts the installer artifacts AwsAccountId: '' AwsBootstrapRoleAccountArn: 'arn:aws:iam:::role/omnistrate-bootstrap-role' ``` | Field | Description | | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | `name` | Display name for your installer product. | | `requirements.k8sVersion` | (Optional) Minimum Kubernetes version required on the target cluster. | | `onPremDeployment.AwsAccountId` | The AWS account ID that hosts the installer artifacts. This is the account where Omnistrate stores and retrieves installer resources. | | `onPremDeployment.AwsBootstrapRoleAccountArn` | The IAM role ARN in the installer artifact hosting account that Omnistrate assumes during installer packaging operations. | ### Helper user script You can use the `onPremInstallerTools.helperUserScript` field to inject common reusable bash functions into the installer artifact. These functions are then available to all action hooks (validate, pre-install, post-install, backup), avoiding duplication across hooks: ``` onPremInstallerTools: helperUserScript: | #!/bin/bash log_error() { echo "Error: $1" > /tmp/error.log } ``` Or reference an external file: ``` onPremInstallerTools: helperUserScript: | {{ $file:./custom_scripts/helper.sh }} ``` ## Defining services Under the `services` key, define one or more services. Each service maps to a deployable unit (typically a Helm release). ``` services: - name: MyApp ``` ## Configuring API parameters API parameters are the user-facing configuration inputs shown during deployment. Add them under `apiParameters` for each service. For more details, see [Deployment API Params](https://docs.omnistrate.com/build-guides/api-params/index.md). ``` apiParameters: - name: releaseName key: releaseName type: String required: false modifiable: false export: true description: "Helm release name" defaultValue: "my-app" - name: namespace key: namespace type: String required: false modifiable: false export: true defaultValue: "my-app" description: "Kubernetes namespace to deploy into" - name: adminPassword key: adminPassword type: Password required: true modifiable: false export: true description: "Admin password for the application" ``` ### Parameter field reference | Field | Description | | -------------- | -------------------------------------------------------------- | | `name` | Human-readable display name. | | `key` | Internal reference key (used in `{{ $var. }}` templates). | | `type` | Data type: `String`, `Boolean`, `Password`, etc. | | `required` | Whether the user must provide a value. | | `modifiable` | Whether the value can be changed after initial deployment. | | `export` | Whether the value is passed to Helm and hooks. | | `defaultValue` | Pre-filled default (user can override). | | `description` | Tooltip or help text shown in the UI. | | `options` | (Optional) Restrict input to an enumerated list of values. | Tip Use `type: Password` for secrets so the value is masked in the UI. ## Configuring action hooks Action hooks run shell commands at specific points in the deployment lifecycle. Add them under `actionHooks`. For more details, see [Action Hooks](https://docs.omnistrate.com/build-guides/actionhooks/index.md). ``` actionHooks: - scope: CLUSTER type: VALIDATE commandTemplate: "echo 'Running validation checks...'" - scope: CLUSTER type: PRE_INSTALL commandTemplate: "echo 'Preparing environment...'" - scope: CLUSTER type: POST_INSTALL commandTemplate: "echo 'Post-install configuration...'" - scope: CLUSTER type: BACKUP commandTemplate: "echo 'Creating backup...'" ``` | Hook Type | When it runs | | -------------- | ----------------------------------------------------------------------- | | `VALIDATE` | Before and after installation to verify the cluster meets requirements. | | `PRE_INSTALL` | After validation but before Helm install/upgrade. | | `POST_INSTALL` | After Helm install/upgrade completes successfully. | | `BACKUP` | Before upgrades to snapshot the current state. | For longer scripts, reference external files: ``` - scope: CLUSTER type: VALIDATE commandTemplate: | {{ $file:./custom_scripts/validate.sh }} ``` ## Configuring the Helm chart The `helmChartConfiguration` block tells Omnistrate which Helm chart to deploy and how to pass values: ``` helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-repo chartRepoURL: https://charts.example.com/ releaseName: "{{ $var.releaseName }}" namespace: "{{ $var.namespace }}" ``` | Field | Description | | --------------- | --------------------------------------------------------------------------- | | `chartName` | Name of the Helm chart. | | `chartVersion` | Chart version to install. | | `chartRepoName` | Logical name for the Helm repository. | | `chartRepoURL` | URL of the Helm chart repository (HTTPS or OCI). | | `releaseName` | Helm release name (use `{{ $var.releaseName }}` to reference a parameter). | | `namespace` | Kubernetes namespace (use `{{ $var.namespace }}` to reference a parameter). | ### Passing values with chartValues Map API parameters into your Helm chart's `values.yaml` structure using the `{{ $var. }}` template syntax: ``` chartValues: config: adminPassword: "{{ $var.adminPassword }}" domain: "{{ $var.dnsEndpoint }}" ingress: enabled: true hosts: - host: "{{ $var.dnsEndpoint }}" paths: - path: / pathType: Prefix ``` ### Platform-specific values with layeredChartValues Use `layeredChartValues` to conditionally apply different Helm values based on the target platform or other parameter values. For more details, see [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md). ``` layeredChartValues: - scope: "{{ $var.onprem_platform }}": "EKS" values: {{ $file:./chart_value_templates/aws/values.yaml }} - scope: "{{ $var.onprem_platform }}": "AKS" values: {{ $file:./chart_value_templates/azure/values.yaml }} - scope: "{{ $var.onprem_platform }}": "GKE" values: {{ $file:./chart_value_templates/gcp/values.yaml }} - scope: "{{ $var.onprem_platform }}": "GENERIC" values: {{ $file:./chart_value_templates/generic/values.yaml }} ``` The `scope` block acts as a conditional — the values are only applied when all conditions match. `{{ $var.onprem_platform }}` is a built-in system variable set automatically based on the customer's cluster type. ## Configuring container image registry copy If your Helm chart references container images stored in a private registry (for example, Docker Hub with authentication, a private ECR, or any registry requiring credentials), the installer needs to know how to pull those images at build time so they can be copied into the customer's target environment. Without this configuration, the customer's cluster may fail to pull images if it has no direct access to your source registry. The `containerImagesRegistryCopyConfiguration` block defines **where to pull images from** (source registry + credentials) and **where to push them to** (the customer's private registry). This is essential for: - **Air-gapped or restricted environments** — the target cluster has no internet access - **Private source registries** — your images require authentication to pull - **Registry locality** — customers want images in their own registry for performance, compliance, or security reasons ### How it works The installer uses a **two-service pattern**: a dedicated image sync service handles the registry copy, and your main application service depends on it. This ensures all images are available in the customer's registry before the Helm chart is installed. ``` ┌─────────────────────┐ ┌─────────────────────┐ │ ImageSync Service │──────▶│ MyApp Service │ │ (internal: true) │ │ (dependsOn: │ │ │ │ - ImageSync) │ │ Copies images from │ │ │ │ source → target │ │ Installs Helm chart│ │ registry │ │ (images now local) │ └─────────────────────┘ └─────────────────────┘ ``` ### Image discovery with autoDiscoverImagesTag The installer needs to know **which images** to copy. Rather than listing every image manually, you can optionally add `autoDiscoverImagesTag` to your `helmChartConfiguration`. This tells the installer to look for the matching annotation in the Helm chart's root `Chart.yaml` metadata and automatically discover all container images referenced by the chart. **In the installer spec**, set the tag value: ``` helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-repo chartRepoURL: oci://registry-1.docker.io/my-org # ...other fields... autoDiscoverImagesTag: "my-org.com/images" ``` **In your Helm chart's `Chart.yaml`**, add the corresponding annotation. Each entry includes a `name` and an `image` reference. You can use Helm's `{{.Chart.AppVersion}}` template to keep image tags in sync with the chart version: ``` apiVersion: v2 name: my-app version: 1.0.0 appVersion: 2.5.0 annotations: my-org.com/images: | - name: app-server image: docker.io/my-org/app-server:{{.Chart.AppVersion}} - name: app-worker image: docker.io/my-org/app-worker:{{.Chart.AppVersion}} - name: mongodb image: docker.io/my-org/mirror-mongodb:7.0-build-335 - name: postgresql-repmgr image: docker.io/my-org/mirror-postgresql-repmgr:14.15.0-debian-12-r2 - name: redis image: docker.io/my-org/mirror-redis-server:7.4-build-248 - name: rabbitmq image: docker.io/my-org/mirror-rabbitmq:4.0.5-build-175 ``` At build time, the installer reads the `Chart.yaml` annotations, matches the `autoDiscoverImagesTag` key, and extracts the listed container images for the registry copy operation. This keeps the image list in sync with your chart — update `Chart.yaml` when you change images, and the installer picks up the changes automatically. Warning Using `autoDiscoverImagesTag` requires a properly configured image registry copy resource (`containerImagesRegistryCopyConfiguration`) in your spec. The auto-discovery feature reads image references from the chart metadata and feeds them into the registry copy pipeline — without a configured image sync service, there is nowhere to copy the discovered images. See [Defining the image sync service](#defining-the-image-sync-service) for setup details. Tip The `autoDiscoverImagesTag` field is optional. If you omit it, you can manually specify the images to copy using the `images` list in `containerImagesRegistryCopyConfiguration` instead (see the [end-to-end example](#end-to-end-example)). Note If your Helm chart is in a private OCI registry, you also need `authProvider` to authenticate when pulling the chart itself: ``` authProvider: username: '{{ $secret.REGISTRY_USERNAME }}' password: '{{ $secret.REGISTRY_PASSWORD }}' ``` ### Pull modes The `pullMode` field controls **when and how** images are transferred: | Mode | Behavior | Best for | | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | `INSTALLER_EMBED` | Images are downloaded from the source registry at **build time** and packaged into the installer artifact. At install time, images are pushed from the local artifact to the target registry. No internet access required during installation. | Air-gapped environments, large-scale deployments where you want deterministic installs. Requires more disk space in the installer artifact. | | `RUNTIME_PULL` | Images are pulled from the source registry and pushed to the target registry at **install/upgrade time**. The installer environment must have network access to both registries. | Connected environments where disk space is a concern, or when you always want the latest image digests. | ### Defining the image sync service Create a separate service with `internal: true` (hidden from end users) that owns the `containerImagesRegistryCopyConfiguration`: ``` services: - name: ImageSync internal: true apiParameters: - name: Private Image Registry URL key: privateRegistryUrl description: "Customer's private registry where images will be pushed" type: String required: true export: true modifiable: true - name: Pull Mode key: pullMode description: "Image sync strategy" type: String required: true export: true modifiable: true defaultValue: INSTALLER_EMBED options: - INSTALLER_EMBED - RUNTIME_PULL containerImagesRegistryCopyConfiguration: pullMode: "{{ $var.pullMode }}" pullSource: registryURL: "docker.io" repositoryName: "my-org" credentials: username: "{{ $secret.REGISTRY_USERNAME }}" password: "{{ $secret.REGISTRY_PASSWORD }}" pushTarget: registryURL: "{{ $var.privateRegistryUrl }}" repositoryName: "my-org" ``` ### Configuration reference | Field | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pullMode` | `INSTALLER_EMBED` or `RUNTIME_PULL` (see [Pull modes](#pull-modes) above). | | `pullSource.registryURL` | Source registry to pull images from (e.g., `docker.io`, `ghcr.io`). | | `pullSource.repositoryName` | Repository or organization name in the source registry (e.g., `my-org`). | | `pullSource.credentials` | (Optional) Credentials for authenticating to the source registry. Use `{{ $secret. }}` to reference Omnistrate-managed secrets. Required if the source registry is private. | | `pushTarget.registryURL` | Customer's target private registry. Typically references an API parameter (e.g., `{{ $var.privateRegistryUrl }}`). | | `pushTarget.repositoryName` | Repository or organization name in the target registry. Usually mirrors the source. | ### Wiring the main service with dependsOn Your main application service must declare `dependsOn` to ensure image sync completes first. Additionally, if the main service needs to reference the same parameters (e.g., `privateRegistryUrl`, `pullMode`), use `parameterDependencyMap` to link them across services rather than duplicating values. For more details on dependencies, see [Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md). ``` - name: MyApp dependsOn: - ImageSync apiParameters: - name: Private Image Registry URL key: privateRegistryUrl description: "Private registry URL" type: String required: true export: true modifiable: true parameterDependencyMap: ImageSync: privateRegistryUrl # passes this value to ImageSync's privateRegistryUrl - name: Pull Mode key: pullMode type: String required: true export: true defaultValue: INSTALLER_EMBED options: - INSTALLER_EMBED - RUNTIME_PULL parameterDependencyMap: ImageSync: pullMode # passes this value to ImageSync's pullMode helmChartConfiguration: chartName: my-app # ... autoDiscoverImagesTag: "my-org.com/images" ``` With `parameterDependencyMap`, the user only fills in the value once (on the main service), and it is automatically forwarded to the internal `ImageSync` service. The format is: ``` parameterDependencyMap: : ``` ### Disabling image sync You can make image sync optional by adding a boolean parameter and using the `disable` field on the image sync service: ``` - name: ImageSync internal: true disable: "{{ $var.skipImageSync }}" apiParameters: - name: Skip Image Sync key: skipImageSync type: Boolean required: false export: true defaultValue: 'false' modifiable: false # ...other parameters... containerImagesRegistryCopyConfiguration: # ... ``` When `skipImageSync` is set to `true`, the entire image copy step is skipped. Use `parameterDependencyMap` on the main service to let the user control this toggle: ``` # In the main service's apiParameters: - name: Skip Image Sync key: skipImageSync type: Boolean required: false export: true defaultValue: 'false' modifiable: false parameterDependencyMap: ImageSync: skipImageSync ``` ### End-to-end example Here is a complete two-service setup with image registry copy: ``` services: # 1. Internal image sync service - name: DockerIO internal: true apiParameters: - name: Skip Custom Image Registry key: skipCustomImageRegistry type: Boolean required: false export: true defaultValue: 'true' modifiable: false - name: Private Image Registry URL key: privateRegistryUrl type: String required: true export: true modifiable: true - name: Pull Mode key: pullMode type: String required: true export: true modifiable: true defaultValue: INSTALLER_EMBED options: - INSTALLER_EMBED - RUNTIME_PULL disable: "{{ $var.skipCustomImageRegistry }}" containerImagesRegistryCopyConfiguration: pullMode: "{{ $var.pullMode }}" pullSource: registryURL: "docker.io" repositoryName: "my-org" credentials: username: "{{ $secret.DOCKERHUB_USERNAME }}" password: "{{ $secret.DOCKERHUB_PASSWORD }}" pushTarget: registryURL: "{{ $var.privateRegistryUrl }}" repositoryName: "my-org" images: - imageName: "my-image" imageTag: "my-tag" # 2. Main application service - name: MyApp dependsOn: - DockerIO apiParameters: - name: Skip Custom Image Registry key: skipCustomImageRegistry type: Boolean required: false export: true defaultValue: 'false' modifiable: false parameterDependencyMap: DockerIO: skipCustomImageRegistry - name: Private Image Registry URL key: privateRegistryUrl type: String required: true export: true modifiable: true parameterDependencyMap: DockerIO: privateRegistryUrl - name: Pull Mode key: pullMode type: String required: true export: true defaultValue: INSTALLER_EMBED options: - INSTALLER_EMBED - RUNTIME_PULL parameterDependencyMap: DockerIO: pullMode # ...other app parameters... helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-org chartRepoURL: oci://registry-1.docker.io/my-org releaseName: "{{ $var.releaseName }}" namespace: "{{ $var.namespace }}" autoDiscoverImagesTag: "my-org.com/images" authProvider: username: '{{ $secret.DOCKERHUB_USERNAME }}' password: '{{ $secret.DOCKERHUB_PASSWORD }}' # ...chartValues, layeredChartValues, etc. ``` ## Advanced: Multiple Helm releases and registries Complex installers often need to install prerequisite Helm charts before the main application. For example, you may want the installer to sync images, install `ingress-nginx`, install `cert-manager`, and then install your application chart. Model this as a graph of services: - Use one internal image sync service for each source registry or repository copy path. - Use one Helm service for each Helm release. - Use `dependsOn` to order image sync before the Helm release that consumes those images. - Use `parameterDependencyMap` to pass shared inputs, such as `privateRegistryUrl`, `pullMode`, and skip flags, from the main application service to internal services. - Use `disable` when an API parameter can decide that a resource should not be included in the installer plan. - Use a validation hook with runtime skip when the decision depends on the customer's cluster state, such as detecting that a prerequisite Helm release already exists. ### Example service graph ``` DockerHubImages -----> MainApp RegistryK8sImages --> IngressNginx --> MainApp QuayJetstackImages --> CertManager --> MainApp ``` The image sync services are hidden from the user with `internal: true`. The prerequisite Helm services can also be internal when they are implementation details of the installer. The examples below use one `prerequisiteAlreadyInstalled` flag to keep the snippets compact. In a production spec, use separate flags when prerequisites can be present independently. ### Shared parameters YAML anchors are optional, but they keep multi-service installer specs readable: ``` x-api-skip-private-registry: &api-skip-private-registry name: Skip Private Image Registry key: skipPrivateRegistry description: Skip copying images into a private registry type: Boolean required: false export: true defaultValue: 'false' modifiable: false x-api-private-registry-url: &api-private-registry-url name: Private Image Registry URL key: privateRegistryUrl description: Private registry where installer images are pushed type: String required: true export: true modifiable: true x-api-pull-mode: &api-pull-mode name: Pull Mode key: pullMode description: Image sync strategy type: String required: true export: true modifiable: true defaultValue: INSTALLER_EMBED options: - INSTALLER_EMBED - RUNTIME_PULL x-api-prereq-exists: &api-prereq-exists name: Prerequisite Already Installed key: prerequisiteAlreadyInstalled description: Skip prerequisite image sync and Helm install when the prerequisite already exists type: Boolean required: false export: true defaultValue: 'false' modifiable: true ``` ### Multiple image sync services Create one registry copy service per source registry or repository layout. This lets each registry keep its own `pullSource`, credentials, image list, and target repository. ``` services: - name: DockerHubImages internal: true apiParameters: - <<: *api-skip-private-registry - <<: *api-private-registry-url - <<: *api-pull-mode disable: "{{ $var.skipPrivateRegistry }}" containerImagesRegistryCopyConfiguration: pullMode: "{{ $var.pullMode }}" pullSource: registryURL: "docker.io" repositoryName: "my-org" credentials: username: "{{ $secret.DOCKERHUB_USERNAME }}" password: "{{ $secret.DOCKERHUB_PASSWORD }}" pushTarget: registryURL: "{{ $var.privateRegistryUrl }}" repositoryName: "my-org" - name: RegistryK8sImages internal: true apiParameters: - <<: *api-prereq-exists - <<: *api-skip-private-registry - <<: *api-private-registry-url - <<: *api-pull-mode disable: '{{ $func.if($func.equals($var.skipPrivateRegistry, "true"), "true", $var.prerequisiteAlreadyInstalled) }}' containerImagesRegistryCopyConfiguration: pullMode: "{{ $var.pullMode }}" pullSource: registryURL: "registry.k8s.io" repositoryName: "ingress-nginx" pushTarget: registryURL: "{{ $var.privateRegistryUrl }}" repositoryName: "ingress-nginx" images: - imageName: "controller" imageTag: "v1.15.1" - imageName: "kube-webhook-certgen" imageTag: "v1.6.9" - name: QuayJetstackImages internal: true apiParameters: - <<: *api-prereq-exists - <<: *api-skip-private-registry - <<: *api-private-registry-url - <<: *api-pull-mode disable: '{{ $func.if($func.equals($var.skipPrivateRegistry, "true"), "true", $var.prerequisiteAlreadyInstalled) }}' containerImagesRegistryCopyConfiguration: pullMode: "{{ $var.pullMode }}" pullSource: registryURL: "quay.io" repositoryName: "jetstack" pushTarget: registryURL: "{{ $var.privateRegistryUrl }}" repositoryName: "jetstack" images: - imageName: "cert-manager-controller" imageTag: "v1.21.0" - imageName: "cert-manager-webhook" imageTag: "v1.21.0" - imageName: "cert-manager-cainjector" imageTag: "v1.21.0" - imageName: "cert-manager-startupapicheck" imageTag: "v1.21.0" ``` ### Multiple Helm services Each Helm release gets its own service and its own `helmChartConfiguration`. Use `dependsOn` to make each chart wait for the image sync service it needs. ``` - name: IngressNginx internal: true dependsOn: - RegistryK8sImages apiParameters: - <<: *api-prereq-exists parameterDependencyMap: RegistryK8sImages: prerequisiteAlreadyInstalled - <<: *api-skip-private-registry parameterDependencyMap: RegistryK8sImages: skipPrivateRegistry - <<: *api-private-registry-url parameterDependencyMap: RegistryK8sImages: privateRegistryUrl - <<: *api-pull-mode parameterDependencyMap: RegistryK8sImages: pullMode disable: "{{ $var.prerequisiteAlreadyInstalled }}" actionHooks: - scope: CLUSTER type: VALIDATE commandTemplate: | export INTERNAL_HELM_PHASE=validate {{ $file:./custom_scripts/ingress_nginx_runtime_skip.sh }} - scope: CLUSTER type: POST_INSTALL commandTemplate: | export INTERNAL_HELM_PHASE=post-install {{ $file:./custom_scripts/ingress_nginx_runtime_skip.sh }} helmChartConfiguration: chartName: ingress-nginx chartVersion: 4.15.1 chartRepoName: ingress-nginx chartRepoURL: https://kubernetes.github.io/ingress-nginx releaseName: ingress-nginx namespace: ingress-nginx runtimeConfiguration: wait: true waitForJobs: true layeredChartValues: - scope: "{{ $var.skipPrivateRegistry }}": 'true' values: global: image: registry: registry.k8s.io - scope: "{{ $var.skipPrivateRegistry }}": 'false' values: global: image: registry: "{{ $var.privateRegistryUrl }}" - name: CertManager internal: true dependsOn: - QuayJetstackImages apiParameters: - <<: *api-prereq-exists parameterDependencyMap: QuayJetstackImages: prerequisiteAlreadyInstalled - <<: *api-skip-private-registry parameterDependencyMap: QuayJetstackImages: skipPrivateRegistry - <<: *api-private-registry-url parameterDependencyMap: QuayJetstackImages: privateRegistryUrl - <<: *api-pull-mode parameterDependencyMap: QuayJetstackImages: pullMode disable: "{{ $var.prerequisiteAlreadyInstalled }}" actionHooks: - scope: CLUSTER type: VALIDATE commandTemplate: | export INTERNAL_HELM_PHASE=validate {{ $file:./custom_scripts/cert_manager_runtime_skip.sh }} - scope: CLUSTER type: POST_INSTALL commandTemplate: | export INTERNAL_HELM_PHASE=post-install {{ $file:./custom_scripts/cert_manager_runtime_skip.sh }} helmChartConfiguration: chartName: cert-manager chartVersion: v1.21.0 chartRepoName: jetstack chartRepoURL: https://charts.jetstack.io releaseName: cert-manager namespace: cert-manager runtimeConfiguration: wait: true waitForJobs: true layeredChartValues: - scope: "{{ $var.skipPrivateRegistry }}": 'true' values: crds: enabled: true imageRegistry: quay.io imageNamespace: jetstack - scope: "{{ $var.skipPrivateRegistry }}": 'false' values: crds: enabled: true imageRegistry: "{{ $var.privateRegistryUrl }}" imageNamespace: jetstack - name: MainApp dependsOn: - DockerHubImages - IngressNginx - CertManager apiParameters: - <<: *api-prereq-exists parameterDependencyMap: RegistryK8sImages: prerequisiteAlreadyInstalled QuayJetstackImages: prerequisiteAlreadyInstalled IngressNginx: prerequisiteAlreadyInstalled CertManager: prerequisiteAlreadyInstalled - <<: *api-skip-private-registry parameterDependencyMap: DockerHubImages: skipPrivateRegistry RegistryK8sImages: skipPrivateRegistry QuayJetstackImages: skipPrivateRegistry IngressNginx: skipPrivateRegistry CertManager: skipPrivateRegistry - <<: *api-private-registry-url parameterDependencyMap: DockerHubImages: privateRegistryUrl RegistryK8sImages: privateRegistryUrl QuayJetstackImages: privateRegistryUrl IngressNginx: privateRegistryUrl CertManager: privateRegistryUrl - <<: *api-pull-mode parameterDependencyMap: DockerHubImages: pullMode RegistryK8sImages: pullMode QuayJetstackImages: pullMode IngressNginx: pullMode CertManager: pullMode helmChartConfiguration: chartName: main-app chartVersion: 1.0.0 chartRepoName: my-org chartRepoURL: oci://registry-1.docker.io/my-org releaseName: main-app namespace: main-app autoDiscoverImagesTag: "my-org.com/images" layeredChartValues: - scope: "{{ $var.skipPrivateRegistry }}": 'true' values: global: imageRegistry: docker.io - scope: "{{ $var.skipPrivateRegistry }}": 'false' values: global: imageRegistry: "{{ $var.privateRegistryUrl }}" ``` ### Runtime skip for existing prerequisites Use `disable` when a user-provided parameter decides whether a resource should run. Use runtime skip when the installer must inspect the target cluster during validation. This is common for prerequisite charts because a customer may already have an ingress controller, cert-manager, CSI driver, or other platform component installed. When a validation hook calls `skip_resource_deployment`, the installer records the skip reason for the current resource and skips its remaining deployment steps during the same run. A retry also uses the persisted skip state. ``` #!/usr/bin/env bash set -uo pipefail skip_current_resource() { local reason="$1" echo "Skipping resource: ${reason}" if declare -F skip_resource_deployment >/dev/null 2>&1; then skip_resource_deployment "$reason" else echo "SKIP: ${reason}" fi exit 0 } helm_release_match() { local pattern="$1" helm list -A --all -o json 2>/dev/null \ | jq -r --arg p "$pattern" ' [.[] | select((.chart // "") | test($p))] | ((map(select(.status == "deployed")) | first) // first // empty) | if type == "object" then "\(.name)|\(.namespace)|\(.chart)|\(.status)" else empty end ' } if [[ "${INTERNAL_HELM_PHASE:-validate}" == "validate" ]]; then match="$(helm_release_match "cert-manager")" if [[ -n "$match" ]]; then skip_current_resource "cert-manager already exists (${match})" fi fi ``` If the hook detects an unhealthy existing installation, exit non-zero instead of skipping. Runtime skip should only be used when the existing resource is acceptable for the main application. ## Building and releasing the installer Use the Omnistrate CLI to build and release the installer: ``` omnistrate-ctl build \ --spec-type ServicePlanSpec \ --file installer-spec.yaml \ --product-name "My Application Air-Gapped Installer" \ --release \ --release-description "v1.0.0 initial release" ``` ## Full Example The following example shows a complete minimal spec that packages Supabase as an air-gapped installer. It includes a deployment block, API parameters for the release name, namespace, JWT secret, database password, and DNS endpoint, along with lifecycle action hooks and a Helm chart configuration. The video walkthrough below demonstrates the end-to-end experience of building, releasing, and installing using this exact spec — [Watch on YouTube](https://www.youtube.com/watch?v=55DbkyemZW4). Complete Supabase specification ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/system-parameters-schema.json name: Supabase deployment: requirements: k8sVersion: ">=1.30.0" # optional onPremDeployment: # AWS account that hosts the installer artifacts AwsAccountId: '' AwsBootstrapRoleAccountArn: 'arn:aws:iam:::role/omnistrate-bootstrap-role' onPremInstallerTools: helperUserScript: | #!/bin/bash log_error() { echo "Error: $1" > /tmp/error.log } services: - name: Supabase apiParameters: - name: releaseName key: releaseName type: String required: false modifiable: false export: true description: "Helm release name" defaultValue: "supabase" - name: namespace key: namespace type: String required: false modifiable: false export: true defaultValue: "supabase" description: "Kubernetes namespace to deploy Supabase into" - name: jwtSecret key: jwtSecret type: Password required: false modifiable: false export: true description: "JWT Secret" - name: dbPassword key: dbPassword description: "Database Password" type: Password modifiable: false export: true required: false - name: dnsEndpoint key: dnsEndpoint type: String required: false modifiable: false export: true description: "DNS Endpoint for Supabase Ingress" defaultValue: "example.com" actionHooks: - scope: CLUSTER type: VALIDATE commandTemplate: "echo 'Validate hook'" - scope: CLUSTER type: PRE_INSTALL commandTemplate: "echo 'Pre-install hook'" - scope: CLUSTER type: POST_INSTALL commandTemplate: "echo 'Post-install hook'" - scope: CLUSTER type: BACKUP commandTemplate: "echo 'Backup hook'" helmChartConfiguration: chartName: supabase chartVersion: 0.1.3 chartRepoName: supabase chartRepoURL: https://example.github.io/supabase-charts/ releaseName: "{{ $var.releaseName }}" namespace: "{{ $var.namespace }}" chartValues: secret: jwt: secret: "{{ $var.jwtSecret }}" db: password: "{{ $var.dbPassword }}" kong: ingress: enabled: true className: "nginx" hosts: - host: "{{ $var.dnsEndpoint }}" paths: - path: / pathType: Prefix ``` ## Template syntax reference | Syntax | Description | | ---------------------------- | ----------------------------------------------------- | | `{{ $var. }}` | References an API parameter by its `key`. | | `{{ $file: }}` | Inlines the contents of a file at build time. | | `{{ $secret. }}` | References an Omnistrate-managed secret. | | `{{ $var.onprem_platform }}` | Built-in variable: `EKS`, `AKS`, `GKE`, or `GENERIC`. | # Air-Gapped Installer Overview ## What Is an Air-Gapped Installer? An air-gapped installer is a self-contained deployment package that allows your customers to install and run your software on their own Kubernetes clusters — including air-gapped environments with no internet connectivity. Instead of managing hosted infrastructure on your customers' behalf, you deliver a ready-to-run installer that packages your Helm chart, container images, configuration, and lifecycle scripts into a single artifact. Air-Gapped is the disconnected end of the BYOC Anywhere spectrum. It is not a live BYOC control-plane connection: the customer owns the connectivity boundary, update boundary, support boundary, and operational evidence boundary. Use it when software updates, container images, licenses, or telemetry cannot assume internet access. This approach is essential for reaching enterprise customers in regulated industries such as finance, healthcare, government, and defense, where strict security policies, data sovereignty requirements, or compliance mandates (HIPAA, GDPR, PCI-DSS, FedRAMP) prevent the use of externally hosted services. For the complete specification reference covering all configuration options, see the [Air-Gapped Installer](https://docs.omnistrate.com/build-guides/air-gapped-helm-charts/index.md) guide. For the broader zero-trust principles behind air-gapped and customer-owned deployments, read [BYOC Anywhere: The 10 Commandments](https://byocanywhere.org/). ### Key Concepts **Installer Artifact** is the packaged output that Omnistrate produces from your Plan specification. It contains everything needed to deploy your application — Helm chart, container images (optionally embedded), configuration templates, and lifecycle scripts. Your customer downloads this artifact and runs it against their cluster. **Plan Specification** is the YAML file that defines your air-gapped installer. It declares deployment requirements, API parameters (user-facing configuration inputs), action hooks (lifecycle scripts), Helm chart configuration, and container image registry copy settings. **Action Hooks** are shell scripts that run at specific points in the deployment lifecycle — validation, pre-install, post-install, and backup. They allow you to enforce prerequisites, prepare the environment, run post-deployment configuration, and create snapshots before upgrades. **Container Image Registry Copy** is the mechanism that transfers your container images from your source registry to the customer's private registry. This is critical for air-gapped deployments where the target cluster has no internet access. ## Offline Delivery Workflow The typical customer workflow is: ``` Receive signed artifacts → scan and approve → import into offline repositories → install or upgrade locally ``` Air-Gapped operations must account for offline updates, mirrored repositories, controlled artifact transfer, signed offline licenses, local diagnostics, and support bundles that do not leak customer data. Do not assume live telemetry, remote debugging, online license checks, automatic image pulls, or continuous control-plane access. For background on disconnected environments, see [Ubuntu's air-gapped system documentation](https://documentation.ubuntu.com/airgapped/about/). ## Why Build an Air-Gapped Installer? ### Reach Regulated and Security-Conscious Customers Many enterprise customers simply cannot use externally hosted services. Data sovereignty laws, internal security policies, and compliance frameworks require software to run within the customer's own infrastructure. An air-gapped installer lets you serve these customers without building a separate version of your product. ### Support Air-Gapped Environments Some of the most demanding deployment scenarios involve networks completely disconnected from the internet. With Omnistrate's `INSTALLER_EMBED` pull mode, container images are bundled directly into the installer artifact, enabling fully offline installations with no external network access required. ### Simplify Customer Deployments Instead of handing customers a collection of Helm charts, scripts, and documentation, you deliver a single installer with a guided deployment experience. API parameters provide a clean configuration interface, and action hooks handle environment preparation and validation automatically. ### Maintain Control and Consistency Even though the software runs on customer infrastructure, you maintain control over the package and deployment process. The installer enforces version requirements, validates cluster prerequisites, and follows a consistent lifecycle — while the customer controls artifact approval, installation, local operations, and evidence within the disconnected environment. ### Streamline Upgrades and Maintenance The installer supports versioned releases with built-in backup hooks, so customers can safely upgrade to new versions. Each new release is a new installer artifact, and the upgrade process includes automatic backup creation before applying changes. ## Common Use Cases for Air-Gapped Installers ### Regulated Industries Deploy into healthcare, financial services, and government environments where compliance frameworks mandate air-gapped or isolated cloud deployments. ### Air-Gapped and Restricted Networks Deliver software to environments with no internet connectivity — military installations, secure government facilities, or isolated industrial networks. ### Customer-Managed Kubernetes Clusters Allow customers to run your software on their existing Kubernetes infrastructure (EKS, AKS, GKE, or bare-metal clusters) without requiring access to your cloud resources. ## Benefits By building air-gapped installers with Omnistrate, you get: - **Enterprise Market Access**: Serve regulated industries and security-conscious customers who require air-gapped deployments - **Air-Gapped Support**: Deliver fully offline installations with embedded container images - **Consistent Deployments**: Enforce version requirements, validate prerequisites, and follow repeatable lifecycle processes - **Multi-Platform Compatibility**: Automatically adapt to EKS, AKS, GKE, and generic Kubernetes clusters - **Simplified Operations**: Package complex multi-component applications into a single, guided installer experience - **Version Management**: Release new installer versions with built-in backup and upgrade workflows - **Single Specification**: Define your installer once in YAML and produce artifacts for all target environments The combination of Helm's packaging capabilities and Omnistrate's installer framework provides a streamlined path from Kubernetes application to enterprise-grade air-gapped distribution. # Deployment API Params ## What are API Params API parameters is a way for you to define the parameters that you want your customers to configure their deployments. Your may have any number of API parameters. You can then configure your service or infrastructure configuration through these API parameters. ## Motivation for API Params One of the biggest complexities with operating SaaS at scale is to meet the needs of different customers with varying requirements. Often, you need a way for your tenants to customize their deployments. The customization could be: - How your application is configured - What capabilities are enabled - What cloud provider, region, account is chosen - What infrastructure is used - How the infrastructure is configured based on the cost and security requirements To enable, we have made sure that you can define and build the experience thats best for your application. To achieve the above, we have defined the concept of parameterization where you can parameterize anything and assign a type to it. ## Types of API Params There are 3 types of parameters: - Static parameters - Dynamic parameters - System parameters We will go over each one of them in detail below. ### Static parameters There are certain parameters that you don't want your customers to worry about and prefer to control them internally. However, you stil want to version it so that when you change them, you can control which tenant gets the updated values or know who has the updated version. ### Dynamic parameters There are certain parameters that you want your customers to configure. You can define such variables and configure them to be dynamic based on the customer input. To achieve this, you can set the value of the corresponding parameters to any API parameter. Now, you may also want to configure infrastructure dynamically based on the customer input. Here are some example infrastructure settings that you can dynamically configure: - replicaCountAPIParam: to configure number of replicas - instanceTypes: to configure instance type - instanceStorageIOPSAPIParam: to configure storage IOPS - instanceStorageThroughputAPIParam: to configure storage throughput - instanceStorageTypeAPIParam: to configure storage type - instanceStorageSizeGiAPIParam: to configure storage size in GBs Note If you need to dynamically customize some specific infra settings, please reach out to us [support@omnistrate.com](mailto:support@omnistrate.com) ## Examples for API Params Here is an example API param for PostgreSQL writer resource. As you can see different API parameter are defined and used to configure database configuration (ex - `postgresqlUsername`) and underlying infrastructure (ex - `writerInstanceType`): ``` services: postgresql: x-omnistrate-api-params: - key: writerInstanceType description: Writer Instance Type name: Writer Instance Type type: String modifiable: true required: true export: true - key: postgresqlPassword description: Default DB Password name: Password type: String modifiable: false required: true export: false regex: ^[a-zA-Z0-9!@#$%^&*()_+]{8,32}$ - key: postgresqlDatabase description: Default DB Name name: Default Database type: String modifiable: false required: true export: true - key: postgresqlUsername description: Username name: Default DB Username type: String modifiable: false required: true export: true regex: ^[a-z0-9_]{8,32}$ - key: postgresqlRootPassword description: Root Password name: Root DB Password type: String modifiable: false required: false export: false defaultValue: rootpassword12345 environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_DATABASE=$var.postgresqlDatabase - POSTGRESQL_USERNAME=$var.postgresqlUsername - POSTGRESQL_POSTGRES_PASSWORD=$var.postgresqlRootPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true - POSTGRESQL_REPLICATION_MODE=master - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=repl_password - POSTGRESQL_DATA_DIR=/var/lib/postgresql/data/dbdata - SECURITY_CONTEXT_USER_ID=1001 - SECURITY_CONTEXT_FS_GROUP=1001 - SECURITY_CONTEXT_GROUP_ID=0 x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: writerInstanceType - cloudProvider: gcp apiParam: writerInstanceType - cloudProvider: azure apiParam: writerInstanceType ``` ## System parameters Finally, there are some system parameters that we have defined in the system that you can use. The values of them are determined at runtime based on other configuration. - Server ID: unique ID of each node in a resource instance. As an example, if you creating Kafka SaaS, then you may define Kafka Cluster as one resource and ZK as another resource. For every instance of Kafka resource, you can use Server ID as a broker ID for each of the brokers (or nodes) - Node name: unique name of each node in a resource instance - Node names: set of all the node names in a given resource instance - Global endpoint: global endpoint of a resource instance that will be load-balanced across all its nodes - Local endpoints: list of local endpoints in a resource instance corresponding to each of the nodes - Bucket name: bucket name if cluster storage is configured In addition, if you define the environment variables SECURITY_CONTEXT_GROUP_ID, SECURITY_CONTEXT_USER_ID and SECURITY_CONTEXT_FS_GROUP, we will automatically use them to configure user ID, group ID and fsGroup ID. Info A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) ## Configuring API Params Each API Parameter can be configured with the following set of properties: - **Key**: key to uniquely identify this API parameter. - **Name**: display name that will be used for your customers to specify the value of this parameter. - **Description**: display description for you to describe this parameter to your customer. - **Type**: data type of the API parameter. We describe the valid [parameters types](#parameter-types) in the next section. - **Required**: specify if this field is a required parameter for your customers. Note that we will require your customers to input the value before they can submit any provisioning request. - **Export**: configures if this field will be returned as part of the describe call on this component. - **Modifiable**: configure if this field is modifiable once configured. - **Default value**: if this parameter is not required, we need a default value for this parameter in case your customer don't input any value to this parameter. *Note that default value is a mandatory field if required parameter is configured as false.* - **Limits**: define lower/upper bounds depending on the type of the parameter. For numbers, its min and max values. For string, its the length of the string. - **Options**: in-case you want your customers to choose between the pre-defined values, you can set this property with a list of possible values. This property is only applicable to numbers, string, JSON types. As an example, you can ask your customer to take `Instance Type` as an API param for your Private GPT SaaS with pre-defined list of GPU instance types. - **Regex**: if you want to validate the input of your customer, you can define a regex that will be used to validate the input. This property is only applicable to string type parameters. - **`tabIndex` (display order)**: integer that controls the display order of this parameter in the Customer Portal and API. Lower values appear first. Parameters without a `tabIndex` value are shown after those with an explicit value. - **Scope**: a set of conditions to determine if this parameter is applicable for a given deployment. For example, you may want to define certain parameters only for AWS deployments. In that case, you can set the scope of that parameter to AWS cloud provider. Currently, only a list of cloud providers is supported as part of scope. ## `tabIndex` Display Order You can control the order in which API parameters appear to your customers in the Customer Portal and API responses by setting the `tabIndex` property. This is useful when you want to present the most important parameters first or group related parameters together. ``` x-omnistrate-api-params: - key: instanceType description: Instance Type name: Instance Type type: String required: true export: true tabIndex: 1 - key: storageSize description: Storage Size (GiB) name: Storage Size type: Float64 required: true export: true tabIndex: 2 - key: enableBackups description: Enable automatic backups name: Enable Backups type: Boolean required: false export: true defaultValue: "true" tabIndex: 3 ``` Parameters are sorted by `tabIndex` in ascending order. Parameters without `tabIndex` appear after all explicitly ordered parameters. ## Parameter types The valid types are: - `boolean`: A true or false value. - `string`: A sequence of characters. - `password`: A sensitive string for passwords, stored securely and masked in outputs. - `float64`: A 64-bit floating-point number. - `bytes`: A sequence of bytes. - `json`: A JSON formatted string. - `any`: Any of the supported types. - `resource`: An API parameter of resource type can be used to link any resource with another resource to enforce creating a linked resource before creating a parent resource. See more details on the [resource linking guide](https://docs.omnistrate.com/build-guides/resource-linking/index.md) ## Dependencies and API Params Now, you may have several resources (or Resource) and they maybe dependent on each other. Let's say if you have resource A (with API param X) that depends on resource B (with API param Y), then you can refer to resource B API and System parameters to configure resource A. More specifically, following variables will be available to configure service/infrastructure parameters of resource A: - Resource A: API parameter X aka $var.X - Resource A: Server ID aka $(NODE_INDEX) - Resource A: Node name aka $sys.compute.node.name - Resource A: Node names aka $sys.compute.nodes[\*].name - Resource A: Global endpoint aka $sys.network.externalClusterEndpoint - Resource A: Local endpoints aka $sys.network.nodes[\*].internalEndpoint - Resource B: API parameter Y aka $B.var.Y - Resource B: Node name aka $B.sys.compute.node.name - Resource B: Node names aka $B.sys.compute.nodes[\*].name - Resource B: Global endpoint aka $B.sys.network.externalClusterEndpoint - Resource B: Local endpoints aka $B.sys.network.nodes[\*].internalEndpoint As an example, let's say you are building Kafka SaaS, and Kafka resource depends on Zookeeper (ZK) resource. You will need to configure Kafka component with ZK endpoint. Note Note that in the above example resource B can't refer to parent resource A. Only parent resources can refer to child resources ## Scoped API Parameters You can define API parameters that are only applicable for specific cloud providers using the `scope` property. This allows you to configure cloud-specific settings that are only required when deploying to certain providers. For example, AWS EFS requires Performance Mode configuration, but this parameter is not relevant for GCP or Azure deployments. By setting the `scope` with specific cloud providers, the parameter will only be displayed and required when customers deploy to those providers. ``` services: exampleResource: x-omnistrate-api-params: - key: performanceMode description: EFS Performance Mode name: Performance Mode type: String modifiable: true required: false export: true defaultValue: "generalPurpose" options: - generalPurpose - maxIO scope: CloudProviders: # Each value in the list is an OR condition - aws ``` ### Current capability boundaries Today, API parameter scoping supports `scope.CloudProviders`, along with the standard parameter controls described above such as type validation, `regex`, `options`, and numeric or string limits. The following behaviors are not currently supported directly in API parameters: - Showing or hiding one parameter based on the value of another parameter - Automatically setting one parameter from another parameter - Advanced JSON-schema-driven UI validation for `json` parameters # BYOC Deployment ## What Is a BYOC Deployment A BYOC (Bring Your Own Cloud) deployment Plan allows your customers to run your software inside their own cloud accounts or Kubernetes clusters while you retain full operational control through the Omnistrate control plane. The Plan specification defines the deployment model, cloud account configuration, API parameters, Helm chart or container configuration, and lifecycle settings. Omnistrate supports three BYOC variants. Each uses the same Plan structure with different deployment configuration. Air-Gapped is a separate disconnected distribution model documented in the [Air-Gapped Deployment Guide](https://docs.omnistrate.com/build-guides/air-gapped-helm-charts/index.md). | Variant | Spec Field | Target | | -------------------- | ----------------------------------------------------- | -------------------------------------------------------------------------------- | | **BYOC-Account** | `byoaDeployment` | Customer cloud accounts (AWS, GCP, Azure, OCI, Nebius) | | **BYOC-VPC** | `byoaDeployment` + customer VPC/network configuration | Existing customer networks with private routes, endpoints, and egress controls | | **BYOC PrivateLink** | `byoaDeployment` + account-level PrivateLink flag | Regulated environments with zero public exposure | | **BYOC-K8s** | `byoaDeployment` | Any customer-managed Kubernetes cluster (cloud-managed, bare-metal, edge, local) | Note The same Plan can serve multiple BYOC variants. The variant is selected at customer account onboarding time, not in the Plan specification itself. ## Architecture references The approved BYOC Anywhere diagrams show the operational boundary for each connected BYOC shape: ### BYOC-Account ### BYOC-VPC ### BYOC-K8s ## Configuring the Deployment Block ### Compose Specification For compose-based specifications, configure your provider account under `x-omnistrate-service-plan`: ``` x-omnistrate-service-plan: name: 'My Product - BYOC' deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role ``` Note Replace the placeholder values with your own cloud account information. You only need to configure the cloud providers you plan to support. For the complete compose specification reference, see [x-omnistrate-service-plan.deployment.byoaDeployment](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-service-plandeploymentbyoadeployment). ### Plan Specification For Plan-based specifications, configure the deployment block at the root level: ``` name: My Product - BYOC deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role ``` For the complete plan specification reference, see [deployment schema](https://docs.omnistrate.com/spec-guides/plan-spec/#deployment-schema). ## Customer Account Onboarding Before deploying into a customer's cloud account, the customer must onboard their account. This establishes the trust relationship between your provisioner account and the customer's account. ### Self-Serve via Customer Portal Customers can onboard their cloud account directly from your [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md). The portal guides them through the cloud-specific setup (CloudFormation for AWS, Cloud Shell for GCP/Azure) and tracks the account status. ### Assisted via CLI You can onboard customer accounts on their behalf using `omnistrate-ctl`: ``` # AWS account omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --aws-account-id= # GCP account omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --gcp-project-id= \ --gcp-project-number= \ --gcp-service-account-email= # Azure account omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --azure-subscription-id= \ --azure-tenant-id= ``` For detailed onboarding guides per cloud provider, see: - [AWS account onboarding](https://docs.omnistrate.com/getting-started/onboarding/aws/index.md) - [GCP account onboarding](https://docs.omnistrate.com/getting-started/onboarding/gcp/index.md) - [Azure account onboarding](https://docs.omnistrate.com/getting-started/onboarding/azure/index.md) - [Nebius account onboarding](https://docs.omnistrate.com/getting-started/onboarding/nebius/index.md) ### Enabling PrivateLink To enable BYOC PrivateLink for a customer account, add the `--private-link` flag at onboarding time: ``` omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --aws-account-id= \ --private-link ``` PrivateLink is a per-account setting — every instance deployed into a PrivateLink-enabled account uses private connectivity automatically. No Plan specification changes are required. For details, see [BYOC PrivateLink](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md). ### Onboarding BYOC-K8s Clusters For BYOC-K8s, pass a cluster name instead of cloud account credentials: ``` mkdir -p dp-install-kit && cd dp-install-kit omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --cluster-name= \ --cluster-description="Customer production Kubernetes cluster" ``` This generates an install kit that the customer runs against their Kubernetes cluster. For the full walkthrough, see [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/#end-to-end-walkthrough). The `byoc-onprem` command values remain unchanged because they are CLI identifiers. ## Bring Your Own VPC (BYO-VPC) Customers running in BYOC mode can bring their own VPC instead of letting Omnistrate create one. When creating an instance, the customer specifies their VPC ID as the value of the `cloud_provider_native_network_id` input parameter. ### VPC Requirements (AWS Standard) | # | Requirement | Details | | --- | -------------------------------- | -------------------------------------------------------------------------------------------------------- | | 1 | **DNS settings** | Enable DNS hostnames and DNS resolution on the VPC | | 2 | **NAT Gateway** | A public NAT gateway for pulling container images; private subnet route tables must route to it | | 3 | **Public subnet auto-assign IP** | Public subnets must have auto-assign public IPv4 address enabled | | 4 | **Subnet tags** | Private subnets: `kubernetes.io/role/internal-elb` = `1`. Public subnets: `kubernetes.io/role/elb` = `1` | ### VPC Requirements (PrivateLink) | # | Requirement | Details | | --- | --------------------------- | ----------------------------------------------------------------------------------------------- | | 1 | **VPC & subnet tags** | Tag VPC and workload subnets with `omnistrate.com/managed-by` = `omnistrate` | | 2 | **DNS settings** | Enable DNS hostnames and DNS resolution | | 3 | **Egress** | Outbound internet via NAT Gateway, Transit Gateway, or VPN | | 4 | **Management VPC Endpoint** | Interface VPC Endpoint targeting the PrivateLink service name Omnistrate provides | | 5 | **Cross-region** | If the customer's VPC and PrivateLink service are in different regions, pass `--service-region` | For full BYO-VPC details, see [BYOC](https://docs.omnistrate.com/usecases/byoc/#bring-your-own-vpc-byo-vpc) and [Imported VPC requirements for BYOC PrivateLink](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#imported-vpc-requirements-for-byoc-privatelink). ## Deploying Instances Once a customer account is onboarded and in `READY` state, deploy instances into it: ``` omnistrate-ctl instance create \ --service= \ --environment= \ --plan= \ --version=latest \ --resource= \ --cloud-provider=aws \ --region=us-east-1 \ --customer-account-id= \ --param-file=./params.json \ --wait ``` For BYOC-K8s deployments, use `--cloud-provider=byoc-onprem` and `--region=on-prem`: ``` omnistrate-ctl instance create \ --service= \ --environment= \ --plan= \ --version=latest \ --resource= \ --cloud-provider=byoc-onprem \ --region=on-prem \ --customer-account-id= \ --param-file=./params.json \ --wait ``` ## What's Next - [BYOC Overview](https://docs.omnistrate.com/build-guides/byoc-overview/index.md) — understand the BYOC model and why it matters - [BYOC Anywhere: The 10 Commandments](https://byocanywhere.org/) — security and operating principles for customer-controlled environments - [BYOC use case](https://docs.omnistrate.com/usecases/byoc/index.md) — architecture and BYO-VPC details - [BYOC PrivateLink](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md) — zero-public-exposure variant - [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) — deploy to any Kubernetes cluster - [BYOC Cloud Accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md) — operational guide for managing customer accounts - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) — compare all deployment models - [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) — full compose spec reference - [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) — full plan spec reference # Build for BYOC Anywhere Overview ## What Is BYOC Anywhere? BYOC Anywhere is a spectrum of deployment and operating models where your software runs inside a customer-controlled environment while you retain operational control through Omnistrate's control plane. Instead of asking customers to manage the software on their own, you provide a managed experience across the customer's account, network, Kubernetes platform, or disconnected environment. This model addresses a fundamental tension in software distribution: customers want the convenience of a managed service, but they need their data, compute, and network traffic to stay within infrastructure they own and control. Omnistrate supports three BYOC variants and a related Air-Gapped distribution model, each targeting a different security and connectivity profile: | Variant | Connectivity | Target Environment | | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | [**BYOC-Account**](https://docs.omnistrate.com/usecases/byoc/index.md) | Customer owns a dedicated account; control traffic uses encrypted channels | Customer cloud accounts (AWS, GCP, Azure) | | [**BYOC-VPC**](https://docs.omnistrate.com/usecases/byoc/index.md) / [**PrivateLink**](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md) | Customer-approved private networking, routes, endpoints, and egress controls; PrivateLink provides zero public exposure for control traffic | Existing customer VPCs, VNets, or private networks, including regulated environments with no-public-egress policies | | [**BYOC-K8s**](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) | Customer provides the Kubernetes runtime; the vendor deploys and operates workloads | Managed Kubernetes, OpenShift, edge, on-premises, or GPU clusters | | [**Air-Gapped**](https://docs.omnistrate.com/build-guides/air-gapped-overview/index.md) | No live control-plane connection; signed artifacts are transferred and installed locally | Disconnected or highly regulated environments | ## Why Build for BYOC? ### Data Sovereignty Is Non-Negotiable Enterprise customers in finance, healthcare, government, and defense operate under strict regulatory frameworks — HIPAA, GDPR, PCI-DSS, SOC 2, FedRAMP — that dictate where data can reside and who can access it. BYOC eliminates the compliance conversation entirely: the data never leaves the customer's infrastructure. ### Security-Conscious Customers Require It Many organizations will not send sensitive data to a third-party cloud account regardless of certifications. BYOC lets you serve these customers without building a separate on-premise version of your product or negotiating complex data processing agreements. ### Cost Efficiency at Scale When customers run workloads in their own accounts, they use their existing cloud commitments — reserved instances, savings plans, enterprise discount programs, and committed-use contracts. This reduces the effective cost of your product and removes the margin pressure of passing through your own infrastructure costs. ### Operational Control Without Infrastructure Ownership Unlike traditional on-premise software where you ship a binary and lose visibility, BYOC through Omnistrate gives you full fleet management capabilities. You can monitor, upgrade, scale, and troubleshoot every deployment across every customer account from a single control plane — as if they were running in your own infrastructure. ## Why BYOC Matters More in the AI Era The rise of AI workloads has made BYOC not just a compliance checkbox but a strategic necessity. ### AI Workloads Amplify Data Sensitivity AI models train on, fine-tune against, and inference over your customers' most valuable data — proprietary datasets, customer records, intellectual property, and trade secrets. The data flowing through AI pipelines is often more sensitive than traditional application data because it reveals business logic, competitive advantages, and strategic direction. Customers will not send this data to your cloud account. ### GPU Costs Demand Customer-Owned Infrastructure AI inference and training require expensive GPU instances. Customers with existing GPU reservations, committed-use discounts, or specialized hardware (NVIDIA H100/H200 clusters, custom FPGA setups) need to run workloads on their own infrastructure to avoid paying retail prices. BYOC lets them leverage their existing compute investments while consuming your product as a managed service. ### Regulatory Frameworks Are Catching Up The EU AI Act, executive orders on AI safety, and emerging frameworks around model governance all push toward stricter controls on where AI workloads run and how training data is handled. BYOC positions your product ahead of these requirements by keeping data and compute within customer-controlled boundaries. ### Edge and On-Premise AI Is Growing Not all AI workloads belong in the cloud. Inference at the edge, on-premise GPU clusters, and air-gapped environments for classified or sensitive workloads are becoming standard deployment targets. BYOC On-Premise supports these scenarios with the same operational model as cloud deployments. ## Build Once, Deploy Anywhere One of Omnistrate's most powerful capabilities is that a single Plan can target any deployment environment. You define your application once and deploy it across the full spectrum of infrastructure: - **Hosted SaaS** — for deployments in your own cloud account - **Customer cloud accounts** — AWS, GCP, Azure via BYOC - **Neoclouds** — Nebius, CoreWeave, Lambda Labs, and other specialized GPU cloud providers - **On-premise Kubernetes** — EKS, AKS, GKE, Rancher, OpenShift, or bare-metal clusters via BYOC On-Premise - **Local development** — k3s, k3d, or minikube for testing and validation - **Air-gapped environments** — fully disconnected networks via the [air-gapped installer](https://docs.omnistrate.com/build-guides/air-gapped-overview/index.md) This means you do not build separate versions of your product for different deployment targets. The same Helm chart, the same configuration, the same lifecycle hooks, and the same operational tooling work everywhere. Your customer chooses the deployment model at subscription time, and Omnistrate handles the rest. ``` graph LR A[Your Plan] --> B[Hosted SaaS] A --> C[BYOC - AWS/GCP/Azure] A --> D[Neoclouds] A --> E[On-Premise K8s] A --> F[Edge / k3s] A --> G[Air-Gapped] ``` ### From k3s to Enterprise Cloud — the Same Control Plane A SaaS provider building an AI inference platform can start by validating their Plan against a local k3s cluster, promote it to their own AWS account for development, deploy it into a customer's GCP account via BYOC, run it on a Nebius GPU cluster for high-performance training workloads, and ship an on-premise installer for a defense customer — all from the same Plan definition and the same Omnistrate control plane. This flexibility is unique to Omnistrate and eliminates the traditional tradeoff between reach and operational complexity. For the broader zero-trust principles behind these architecture models, read [BYOC Anywhere: The 10 Commandments](https://byocanywhere.org/). ## BYOC Variants in Detail ### BYOC-Account [BYOC-Account](https://docs.omnistrate.com/usecases/byoc/index.md) deploys your application into your customer's cloud account (AWS, GCP, or Azure). Omnistrate establishes a trust relationship between your account and the customer's account, then uses secure, encrypted channels with mTLS to manage the deployment lifecycle. Your customer onboards their cloud account through a CloudFormation stack (AWS), Terraform module (GCP/Azure), or the Customer Portal. Once onboarded, you can deploy, monitor, upgrade, and troubleshoot instances in their account as if they were your own. Customers can also bring their own VPC (BYO-VPC) for tighter network control over where workloads are placed. ### BYOC-VPC and PrivateLink BYOC-VPC deploys into a customer-approved network boundary rather than assuming a fresh account or default network. The deployment must respect the customer's VPC or VNet, subnets, routing, DNS, firewall, private endpoints, and egress policies. This variant includes BYOC PrivateLink for environments that require zero public exposure for control traffic. This model is suitable when customers require private connectivity, centralized firewall inspection, egress allowlists, internal DNS and certificate policies, or existing network segmentation. See [Customer Networks](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md) and [BYOC Deployment: Bring Your Own VPC](https://docs.omnistrate.com/build-guides/byoc-deployment/#bring-your-own-vpc-byo-vpc). [BYOC PrivateLink](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md) adds a zero-public-exposure guarantee on top of standard BYOC-VPC. All control-plane traffic between the customer's dataplane and your Omnistrate-managed control plane flows over AWS PrivateLink. The customer's EKS cluster has no public endpoint and no internet-facing load balancers for control traffic. This variant is selected per customer account at onboarding time — no compose-spec changes required. It is essential for customers with strict no-public-egress policies, typically found in regulated financial services and government environments. ### BYOC-K8s [BYOC-K8s](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) extends the BYOC model to any Kubernetes cluster the customer operates — cloud-managed or bare-metal, in a data center or at the edge. The customer provides the runtime and platform standards; you deploy and operate the application through Helm, operators, or other Kubernetes-native mechanisms. The customer may enforce admission policies, image scanning, secrets management, storage classes, ingress controllers, service mesh rules, node pools, GPU scheduling, and observability integrations. Cluster version, CNI behavior, storage drivers, quotas, and registry access therefore become shared responsibilities with the customer's platform team. This variant supports the widest range of deployment targets, from enterprise Kubernetes platforms (EKS, AKS, GKE, OpenShift, Rancher) to lightweight distributions (k3s, k3d) and purpose-built GPU clusters. ## How each model maps to customer needs | Customer need | Best-fit model | | ----------------------------------------------------------------- | ---------------------- | | Use committed cloud spend and keep a clean ownership boundary | BYOC-Account | | Enforce private networking and customer-controlled routes | BYOC-VPC / PrivateLink | | Reuse internal platform standards and approved Kubernetes tooling | BYOC-K8s | | Meet disconnected, classified, or strict sovereignty requirements | Air-Gapped | Across every model, BYOC Anywhere must support the full lifecycle: provision, deploy, configure, govern, upgrade, meter, observe, and operate. This requires least-privilege permissions, zero-inbound access where possible, egress allowlists, supply-chain integration, auditability, and customer-safe observability. ## Getting Started 1. **Build a Plan** using a [compose specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) or [plan specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) that defines your application's resources, parameters, and lifecycle 1. **Configure BYOC deployment** in your Plan specification — see [deployment models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) for the schema 1. **Onboard customer accounts** through the Customer Portal or `omnistrate-ctl` — see [BYOC cloud accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md) for the operational guide 1. **Deploy and operate** instances across your customer fleet from the Omnistrate control plane # Cellular Multi-Tenancy Deployment Guide This guide shows how to configure Cellular Multi-Tenancy for Hosted SaaS deployments. ## Hosted SaaS For Compose-based SaaS Products, set `OMNISTRATE_MULTI_TENANCY` and configure the provider account with `hostedDeployment`: ``` x-omnistrate-service-plan: name: 'Cellular Multi-Tenant SaaS' tenancyType: 'OMNISTRATE_MULTI_TENANCY' deployment: hostedDeployment: awsAccountId: '' awsBootstrapRoleAccountArn: 'arn:aws:iam:::role/omnistrate-bootstrap-role' ``` Configure resource requests and limits so Omnistrate can bin-pack customer deployments safely. For placement controls, see [Deployment Cells](https://docs.omnistrate.com/build-guides/deployment-cells/index.md) and [Custom Deployment Cell Placement](https://docs.omnistrate.com/infra-guides/custom-deployment-cell-placement/index.md). ## What's Next - [Cellular Multi-Tenancy Overview](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-overview/index.md) - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/#multi-tenancy) - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) - [Deployment Cells](https://docs.omnistrate.com/build-guides/deployment-cells/index.md) # Cellular Multi-Tenancy Cellular Multi-Tenancy lets multiple customer deployments share infrastructure while maintaining logical isolation. Customers are grouped into deployment cells so you can balance utilization, fault isolation, and operational scale. ## When to use Cellular Multi-Tenancy This model is suitable for: - Cost-efficient SaaS tiers and high customer density. - Free or entry-level plans with moderate isolation requirements. - Applications that are tenant-aware or can group tenants into cells. - Services that benefit from automated bin-packing and shared infrastructure. Omnistrate manages placement, resource allocation, scaling, and tenant isolation. For the complete tenancy behavior and placement options, see [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/#multi-tenancy). For configuration and rollout steps, see the [Cellular Multi-Tenancy Deployment Guide](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-deployment/index.md). ## Related Guides - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) - [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) - [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) # Compose Deployment Strategy Omnistrate lets you turn an existing Docker Compose application into a SaaS Product without rewriting the application topology. Compose remains the source for services, networks, volumes, and configuration while Omnistrate adds the product and operations layer around it. ## When to Use Compose Compose is a good fit when: - Your application is already defined by a Docker Compose file. - You want to package several containers and their relationships as one deployable product. - You want to expose customer inputs, endpoints, tenancy, observability, billing, and lifecycle operations around the Compose application. - You prefer a concise YAML definition over a Kubernetes-native or infrastructure-as-code model. If your application lifecycle is managed by custom resources, use the [Kubernetes Operators strategy](https://docs.omnistrate.com/build-guides/operators-overview/index.md). If your product requires cloud infrastructure managed as code, use the [Terraform strategy](https://docs.omnistrate.com/build-guides/terraform-overview/index.md). ## How Compose Fits into Omnistrate The Compose file describes the application topology. Omnistrate extensions add the service plan, resource visibility, compute, storage, API parameters, integrations, action hooks, and deployment capabilities needed to deliver that topology as a managed service. The resulting control plane can manage customer subscriptions and deployments across supported clouds and deployment models while the application continues to run from the Compose definition. ## Strategy Guide - Start with [Build from Compose](https://docs.omnistrate.com/getting-started/build-from-compose/index.md) for the short onboarding path. - Use the [Compose Service Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) for the complete Omnistrate Compose extension reference. - See [Compose Troubleshooting](https://docs.omnistrate.com/build-guides/compose-troubleshooting/index.md) when a Compose deployment or workflow fails. - Review the [generic troubleshooting workflow](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md) for shared debugging tools and instance-debug behavior. For the high-level Docker Compose concepts and schema entry point, see the [Compose specification overview](https://docs.omnistrate.com/spec-guides/compose-spec/index.md). # Docker Compose Service Specification ## Docker Compose Service Specification Overview Omnistrate extends the Docker Compose specification, which allows custom extensions using the syntax `x-`, and supports a subset of the standard specifications. Omnistrate leverages the standard spec and those extensions together to complete the spec for your SaaS. We will go over the [extensions](#custom-tags) below in more detail. Note When you use the `build-from-repo` command, Omnistrate automatically generates a Docker Compose file for you if one doesn't already exist in your project. This generated Compose file includes the necessary configuration based on your container image and any environment variables you specify during the build process. ## Using Schema Validation To simplify the definition of your Omnistrate compose specification, a JSON schema is available that provides validation and autocompletion for all Omnistrate custom extensions. You can use the JSON schema in IDEs that use the YAML Language Server (eg: VSCode / NeoVim). Add the following line at the top of your compose file: ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/compose-spec-schema.json ``` For [**IntelliJ**](https://www.jetbrains.com/help/idea/yaml.html#use-schema-keyword) replace the top line with the following line to set up the yaml schema: ``` # $schema: https://api.omnistrate.cloud/2022-09-01-00/schema/compose-spec-schema.json ``` The schema extends the upstream Docker Compose 3.9 JSON schema with typed definitions for all `x-omnistrate-*` extensions. It provides validation for extension fields at the correct levels (top-level, service-level, and volume-level). You can also retrieve the schema programmatically: ``` GET https://api.omnistrate.cloud/2022-09-01-00/json-schema?type=compose ``` ## Docker Compose File Format The Compose file is a YAML file defining: - **Version** (Optional) - **Services** (Required) - **Networks** - **Volumes** - **Configs** - **Secrets** The default path for a Compose file is `omnistrate-compose.yaml`. Omnistrate supports the Docker Compose 3.9 specification. ``` version: '3.9' ``` ### Basic Structure ``` services: web: image: nginx:latest ports: - "80:80" volumes: - ./html:/usr/share/nginx/html environment: - ENV_VAR=value volumes: data: networks: frontend: ``` ## Supported Native Tags Here are the native tags Omnistrate supports either natively or with an optimized implementation: ### image Specifies the image to start the container from. The image must follow the Open Container Specification addressable image format, as `[/][/][:|@]`. ``` services: web: image: redis # or image: redis:5 # or image: redis@sha256:0ed5d5928d4737458944eb604cc8509e245c3e19d02ad83935398bc4b991aac7 # or image: my_private.registry:5000/redis ``` Native support for all public registries and private registries with credentials. Tip Prefer immutable image references such as versioned tags or digests over mutable tags like `latest`, especially for production rollouts and sidecars. Mutable tags make promotions harder to reason about and can cause older cached images to continue running after redeployments. ### expose Defines the (incoming) port or a range of ports that Compose exposes from the container. These ports must be accessible to linked services and should not be published to the host machine. ``` services: web: expose: - "3000" - "8000" - "8080-8085/tcp" ``` ### ports Exposes container ports. Port mapping must not be used with `network_mode: host`. **Short syntax:** ``` services: web: ports: - "3000" - "3000-3005" - "8000:8000" - "9090-9091:8080-8081" - "127.0.0.1:8001:8001" - "6060:6060/udp" ``` **Long syntax:** ``` services: web: ports: - name: web target: 80 protocol: tcp ``` ### volumes Define mount host paths or named volumes that are accessible by service containers. **Short syntax:** ``` services: web: volumes: - /var/lib/mysql - /opt/data:/var/lib/mysql - ./cache:/tmp/cache - ~/configs:/etc/configs/:ro ``` **Long syntax:** ``` services: web: volumes: - type: volume source: mydata target: /data volume: nocopy: true - type: bind source: ./static target: /opt/app/static ``` ### environment Defines environment variables set in the container. Environment variables can use either an array or a map. **Map syntax:** ``` services: web: environment: RACK_ENV: development SHOW: "true" USER_INPUT: ``` **Array syntax:** ``` services: web: environment: - RACK_ENV=development - SHOW=true - USER_INPUT ``` #### env_file Adds environment variables to the container based on the file content. ``` services: web: env_file: .env # or env_file: - ./a.env - ./b.env ``` **Env file format:** Each line in an `.env` file must be in `VAR[=[VAL]]` format. The following syntax rules apply: - Lines beginning with `#` are processed as comments and ignored - Blank lines are ignored - Values can optionally be quoted - Variable interpolation is supported using `${VARIABLE_NAME}` syntax Example `.env` file: ``` # Set Rails/Rack environment RACK_ENV=development VAR="quoted" DB_HOST=${DB_HOST} ``` ### depends_on Expresses startup and shutdown dependencies between services. **Short syntax:** ``` services: web: depends_on: - db - redis redis: image: redis db: image: postgres ``` **Long syntax:** ``` services: web: depends_on: db: condition: service_healthy restart: true redis: condition: service_started ``` ### entrypoint Declares the default entrypoint for the service container. This overrides the `ENTRYPOINT` instruction from the service's Dockerfile. ``` services: web: entrypoint: /code/entrypoint.sh # or entrypoint: - php - -d - zend_extension=/usr/local/lib/php/extensions/no-debug-non-zts-20100525/xdebug.so - -d - memory_limit=-1 - vendor/bin/phpunit ``` ### command Overrides the default command declared by the container image. ``` services: web: command: bundle exec thin -p 3000 # or command: ["bundle", "exec", "thin", "-p", "3000"] ``` ### init Runs an init process (PID 1) inside the container that forwards signals and reaps processes. Tag is ignored and supported through cluster_init action hooks, which runs as a job. ### container_name / hostname These tags are ignored. Omnistrate has a fully managed DNS service that doesn't require any explicit configuration to work. ### networks Networks tags are ignored as Omnistrate auto-configures network and ports as part of the tenancy model. ### tmpfs These tags are ignored. Omnistrate automatically mounts a temp file on /tmp folder. ### logging These tags are ignored. Omnistrate offers its own managed logging system as integration that can be enabled by adding the `x-customer-integrations` custom spec followed by the `logs` item. This tag is ignored and supported through `x-customer-integrations`. ### healthcheck These tags are ignored. Healthchecks are configured via `x-omnistrate-action-hooks`. Declares a check that's run to determine whether or not the service containers are "healthy". This tag is ignored and supported through `x-omnistrate-action-hooks`. ### user User definition is achieved via `SECURITY_CONTEXT` environment variables according to the UID and GID. ``` services: web: user: "1001" # or user: "1001:1001" ``` ### ulimits Overrides the default ulimits for a container. ``` services: web: ulimits: nproc: 65535 nofile: soft: 20000 hard: 40000 ``` ### cap_add / cap_drop Specifies additional container capabilities or capabilities to drop. ``` services: web: cap_add: - ALL cap_drop: - NET_ADMIN - SYS_ADMIN ``` ### sysctls These tags are ignored. Omnistrate sets up default configuration. ### labels Allows you to specify custom metadata for resources, such as custom resource names or descriptions. ``` services: web: labels: - "name=Your Resource Name" - "description=Your Resource Description" # or labels: com.example.description: "Accounting webapp" com.example.department: "Finance" ``` ### platform Defines the target platform the containers for the service run on. It uses the `os[/arch[/variant]]` syntax. ``` services: web: platform: linux/amd64 # or platform: linux/arm64 ``` The supported values by Omnistrate are `linux/amd64` and `linux/arm64`. It is only used when an instance type is not explicitly defined. ### deploy Specifies the configuration for the deployment and lifecycle of services. This includes resource constraints and requirements. ``` services: web: deploy: resources: limits: cpus: '1000m' memory: 256M reservations: cpus: '100m' memory: 100M ``` **Resource configuration:** - **limits**: Define the maximum resources the container can use - **cpus**: Maximum CPU allocation (can be specified as decimal or with 'm' suffix for millicores) - **memory**: Maximum memory allocation (supports units like M, G) - **reservations**: Define the minimum resources guaranteed to the container - **cpus**: Minimum CPU allocation - **memory**: Minimum memory allocation ### privileged Configures the service container to run with elevated privileges. Support and actual impacts are platform specific. ``` services: web: privileged: true ``` ## Custom Tags In addition, Omnistrate also supports several custom tags as follows. ### x-omnistrate-service-plan `x-omnistrate-service-plan` allows you to configure the Plan for your SaaS in the compose specification file. For the Cellular Multi-Tenancy and Dedicated Tenancy use cases, see the [Cellular Multi-Tenancy Deployment Guide](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-deployment/index.md) and [Dedicated Tenancy Deployment Guide](https://docs.omnistrate.com/build-guides/dedicated-tenancy-deployment/index.md). ``` x-omnistrate-service-plan: name: 'Mysql Free Tier' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' features: CUSTOM_NETWORKS: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 1 deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' ``` Configuration fields: - **name**: The name of the Plan - **tenancyType**: The tenancy type - **deployment**: The deployment type of the Plan. Options: - `hostedDeployment`: The deployed instances of the plan are deployed through Hosted SaaS - `byoaDeployment`: The deployed instances of the plan is deployed on your customer's account ### x-omnistrate-service-plan.name Name of the service plan for your service. ``` x-omnistrate-service-plan: name: 'Service Plan Name' ``` ### x-omnistrate-service-plan.tenancyType Define the tenancy type for your service: ``` x-omnistrate-service-plan: tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' ``` **tenancyType**: The tenancy type of the Plan. Options: - `OMNISTRATE_DEDICATED_TENANCY`: Infrastructure dedicated to a single customer deployment - `OMNISTRATE_MULTI_TENANCY`: Infrastructure shared among multiple customers and deployments - `CUSTOM_TENANCY`: Not valid for compose specification, only used by Plan Specification ### x-omnistrate-service-plan.enableDeletionProtection Enable deletion protection for your service plan to prevent accidental deletions of customer instances. ``` x-omnistrate-service-plan: enableDeletionProtection: true ``` ### x-omnistrate-service-plan.deployment.hostedDeployment Configure the accounts for your hosted service Plan. ``` x-omnistrate-service-plan: deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' ``` Note OCI is not supported in Compose specifications. To deploy on OCI, use the [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) with `CUSTOM_TENANCY` tenancy type instead. ### x-omnistrate-service-plan.deployment.byoaDeployment Configure the intermediate account for your control plane that will act as provisioner for the services running on your customer accounts (BYOC). Customers will provide the configured account permissions to manage their services. ``` x-omnistrate-service-plan: deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role ``` ### x-omnistrate-service-plan.features.CUSTOM_DEPLOYMENT_CELL_PLACEMENT The `CUSTOM_DEPLOYMENT_CELL_PLACEMENT` feature allows you to control how many deployments can be co-located on the same host cluster (deployment cell / Kubernetes cluster). This provides fine-grained control over deployment isolation and resource allocation. ``` x-omnistrate-service-plan: name: 'PostgreSQL Service' features: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 1 ``` ### x-omnistrate-service-plan.features.CUSTOM_NETWORKS The `CUSTOM_NETWORKS` feature allows your customers to define network partitioning on a dedicated stack while keeping the service deployed in the Service Provider Account. This provides complete isolation by provisioning a dedicated stack that is not shared with other customers, enabling private network connectivity while maintaining self-service capabilities. ``` x-omnistrate-service-plan: name: 'PostgreSQL Service' features: CUSTOM_NETWORKS: ``` > **Note:** For `deployment`, replace the account numbers, project id and other information with your own account information. ### x-omnistrate-my-account (deprecated) `x-omnistrate-my-account` is deprecated and allowed you to configure your cloud provider account details for Hosted SaaS. Use the `x-omnistrate-service-plan.deployment.hostedDeployment` tag to configure your cloud account. ### x-omnistrate-byoa (deprecated) `x-omnistrate-byoa` is deprecated and allowed you to configure your cloud provider account details for BYOC. Use the `x-omnistrate-service-plan.deployment.byoaDeployment` tag to configure your cloud account. ### x-omnistrate-mode-internal `x-omnistrate-mode-internal` tag allows you to tag any Resource in the compose specification as internal. ``` services: internal-service: image: my-internal-service x-omnistrate-mode-internal: true ``` - Possible values are `true` or `false` - Default value is `false` - If set to `true`, this Resource will be internal and won't be exposed to your customers - If set to `false`, the Resource will be exposed to your customers to configure and provision them ### x-omnistrate-api-params `x-omnistrate-api-params` allows you to define API params in addition to environment variables. ``` services: database: image: postgres x-omnistrate-api-params: - key: writerInstanceType description: Writer Instance Type name: Writer Instance Type type: String modifiable: true required: true export: true ``` Each API parameter can be configured to specify: - **key**: Unique identifier for the parameter - **name**: Display name of the parameter - **description**: Description for the parameter - **type**: Type of the parameter (String, Float64, Boolean, etc.) - **required**: Specify if this field is a required parameter - **export**: Configures if this field will be returned as part of the describe call - **modifiable**: Configure if this field is modifiable once configured - **defaultValue**: Default value for this field - **options**: List of options for the user to choose from - **labeledOptions**: List of options with user-friendly labels - **limits**: Minimum and maximum values for this field - **regex**: Regex pattern to validate the input - **tabIndex**: Integer that controls the display order of this parameter in the Customer Portal. Lower values appear first. Parameters without `tabIndex` are shown after those with an explicit display order. API responses are not guaranteed to be returned in `tabIndex` order unless an endpoint explicitly documents that behavior. - **scope**: A set of conditions to determine if this parameter applies to a given deployment. Use `CloudProviders` to restrict a parameter to specific clouds. See [Scoped API Parameters](https://docs.omnistrate.com/build-guides/api-params/#scoped-api-parameters). **Example with cloud provider scope:** ``` x-omnistrate-api-params: - key: performanceMode description: EFS Performance Mode name: Performance Mode type: String modifiable: true required: false export: true defaultValue: "generalPurpose" options: - generalPurpose - maxIO scope: CloudProviders: - aws ``` This parameter is only displayed when the customer deploys to AWS. For GCP or Azure deployments, the parameter is hidden; in this example, the configured `defaultValue` is applied when the parameter is not shown. **Example with options:** ``` x-omnistrate-api-params: - key: writerInstanceType description: Writer Instance Type name: Writer Instance Type type: String modifiable: true required: true export: true options: - t4g.small - t4g.medium - t4g.large ``` **Example with labeled options:** ``` x-omnistrate-api-params: - key: writerInstanceType description: Writer Instance Type name: Writer Instance Type type: String modifiable: true required: true export: true labeledOptions: Small: t4g.small Medium: t4g.medium Large: t4g.large ``` **Example with limits:** ``` x-omnistrate-api-params: - key: postgresqlRootPassword description: Postgresql Root Password name: Postgresql Root Password type: String modifiable: true required: true export: true limits: minLength: 8 maxLength: 64 - key: postgresqlPort description: Postgresql Port name: Postgresql Port type: Float64 modifiable: true required: true export: true limits: min: 1024 max: 65535 ``` **Example with default value:** ``` x-omnistrate-api-params: - key: postgresqlUsername description: Postgresql Username name: Postgresql Username type: String modifiable: true required: true export: true defaultValue: postgres ``` **Example with regex validation:** ``` x-omnistrate-api-params: - key: postgresqlPassword description: Default DB Password name: Password type: String modifiable: false required: true export: false regex: ^[a-zA-Z0-9!@#$%^&*()_+]{8,32}$ ``` API Parameter Types: The valid parameter types are: - **boolean**: A true or false value - **string**: A sequence of characters - **password**: A sensitive string for passwords, stored securely and masked in outputs - **float64**: A 64-bit floating-point number - **bytes**: A sequence of bytes - **json**: A JSON formatted string - **any**: Any of the supported types - **resource**: Used for resource linking to enforce creating a linked resource before creating a parent resource An API parameter of resource type can be used to link any resource with another resource to enforce creating a linked resource before creating a parent resource. ``` services: Proxy: image: omnistrate/pgadmin4:7.5 x-omnistrate-api-params: - key: database description: backend database name: database type: Resource export: false required: true modifiable: false Database: image: 'bitnami/postgresql:latest' ``` In this example, customers must create a Database instance first, then create a Proxy instance by specifying which Database instance to link to. You can map API parameters between dependent resources: ``` services: Cluster: image: omnistrate/noop depends_on: - Writer - Reader x-omnistrate-api-params: - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: true export: true defaultValue: t4g.small parameterDependencyMap: Writer: writerInstanceType Reader: readerInstanceType ``` This maps the `instanceType` parameter of the Cluster resource to `writerInstanceType` of the Writer resource and `readerInstanceType` of the Reader resource. > **Note:** All environment variables are **automatically** added as API parameters if `x-omnistrate-api-params` is not specified. ### x-omnistrate-actionhooks `x-omnistrate-actionhooks` allows you to configure action hooks for a given Resource. ``` services: database: image: postgres x-omnistrate-actionhooks: - scope: CLUSTER type: INIT commandTemplate: > PGPASSWORD={{ $var.postgresqlRootPassword }} psql -U postgres -h writer {{ $var.postgresqlDatabase }} -c "create extension vector" ``` Action hooks allow you to inject custom code at different phases of the lifecycle of your control plane operations. In the above example, we are enabling vector extension for Postgres on cluster initialization. Action hooks are categorized into two scopes: **Node Scope** - runs for every node of the Resource: - **`HEALTH_CHECK`**: Periodic health checks mapped to Kubernetes liveness probes - **`READINESS_CHECK`**: Determines service readiness for traffic (defaults to `HEALTH_CHECK`) - **`STARTUP_CHECK`**: Ensures service startup completion (defaults to `HEALTH_CHECK`) - **`POST_START`**: Executes after node starts - **`ADD`**: Triggers when a new node is added - **`REMOVE`**: Triggers when a node is removed - **`INIT`**: Runs initialization before node starts - **`PROMOTE`**: Triggers when a node is promoted to primary role - **`DEMOTE`**: Triggers when a node is demoted before replacement **Cluster Scope** - runs for the entire Resource: - **`POST_START`**: Executes after all nodes have started - **`PRE_START`**: Runs before all nodes are started - **`POST_STOP`**: Executes after all nodes are stopped - **`PRE_STOP`**: Runs before all nodes are stopped - **`POST_UPGRADE`**: Executes after all nodes are upgraded - **`PRE_UPGRADE`**: Runs before all nodes are upgraded - **`INIT`**: Runs initialization before all nodes are started Action hooks support additional configuration options: ``` services: high-performance-db: image: postgres:15 x-omnistrate-actionhooks: - image: busybox:1.37 scope: NODE type: INIT command: - /bin/sh - -c commandTemplate: | # Set the AIO max number of events sysctl -w fs.aio-max-nr='102052' ``` **Configuration fields:** - **image**: Custom container image for the action hook - **command**: Command array to execute - **commandTemplate**: Template for the command with variable interpolation - **scope**: NODE or CLUSTER - **type**: Action hook type as listed above ### x-omnistrate-compute `x-omnistrate-compute` allows you to customize the compute parameters of a Resource. You can customize the following: - Number of replicas - Instance type across cloud providers - Size of the root volume (GiB) across instances ``` services: web: image: nginx x-omnistrate-compute: replicaCountAPIParam: numReplicas instanceTypes: - cloudProvider: aws name: t4g.small - cloudProvider: gcp name: e2-medium - cloudProvider: azure name: Standard_B2als_v2 rootVolumeSizeGi: 10 ``` In the above example, `replicaCountAPIParam` is configured dynamically using `numReplicas` parameter. The users of your service can decide how many replicas they want. Similarly, `instanceTypes` can also be dynamic if you want your customers to specify the machine type among the list of options. For `rootVolumeSizeGi`, you can specify any integer between 10 and 16384 (up to 16 TB). You can use the `instanceTypes` list to override the compute instance type per cloud provider. These explicit instance type entries take precedence over `platform`-based defaults for the provider where they are defined. For example: ``` services: web: image: ghcr.io/example/web:latest platform: linux/arm64 x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: r7g.large # ARM64 - cloudProvider: gcp name: n2-highmem-2 # x86_64 ``` When you override instance types across clouds, ensure the selected machine architecture is compatible with the container image you provide. Use a multi-architecture container image if you intentionally mix ARM64 and x86_64 instance types. Instance type overrides are resource-local and are not inherited through `depends_on`. In a large resource tree, define `instanceTypes` only on compute-backed resources that need the override. To reuse the same selected values across multiple resources, use provider-specific `apiParam` values with `parameterDependencyMap`, and make each compute-backed resource reference those parameters in its own `x-omnistrate-compute.instanceTypes` list. ### x-omnistrate-compute.instanceTypes.configurationOverrides.acceleratorConfiguration GPU accelerator configuration enables you to specify dedicated GPU resources and is only required to attach for some instance types on GCP. 1. **External GPU attachment** using `acceleratorConfiguration` (N1 instances + Tesla GPUs only) 1. **Built-in GPU instances** (G2, A2, A3, A4 families with modern GPUs) The `acceleratorConfiguration` feature allows declarative specification of GPU accelerators attached to compute instances on GCP. > **Warning:** External GPU attachment is **only supported** with N1 instances and Tesla series GPUs (T4, V100, P100, P4) on GCP **Configuration in x-omnistrate-compute:** ``` services: gpu-service: image: tensorflow/tensorflow:latest-gpu x-omnistrate-compute: instanceTypes: - name: n1-standard-8 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-t4" count: 1 ``` **Multi-GPU Configuration:** ``` services: gpu-cluster: image: pytorch/pytorch:latest x-omnistrate-compute: instanceTypes: - name: n1-standard-16 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-t4" count: 4 ``` For AWS, Azure and modern GPUs in GCP (L4, A100, H100) , instance families come with GPUs mounted: **G2 Family - NVIDIA L4 GPUs:** ``` services: modern-gpu-service: image: nvidia/cuda:12.0-runtime-ubuntu20.04 x-omnistrate-compute: instanceTypes: - name: g2-standard-8 cloudProvider: gcp # No acceleratorConfiguration needed - L4 GPU is built-in ``` ### x-omnistrate-storage To define a persistent volumes that can survive service restarts and nodes replacements you can use `x-omnistrate-storage` tag. You can optionally define the volume: ``` volumes: pg_master_data: driver: local services: database: image: postgres volumes: - source: pg_master_data target: /var/lib/postgresql/data type: volume x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP3 instanceStorageSizeGi: 10 instanceStorageIOPS: 3000 instanceStorageThroughputMiBps: 125 gcp: instanceStorageType: GCP::PD_BALANCED instanceStorageSizeGi: 100 azure: instanceStorageType: AZURE::PREMIUM_SSD instanceStorageSizeGi: 100 ``` `x-omnistrate-storage` allows you to customize the storage parameters of a Resource. You can customize the following: - Storage type from block device to blobs or both - Size of the volume - Storage IOPS - Storage throughput Note This type of persistent volume are not shared across Resources. If you need to share the volume data across resources look a at configuring Shared File System or Blob Storage instead. ### volumes.driver: sharedFileSystem Shared file systems allow you to save data in a persistent volume and share it across different pods within the same deployment cluster. This is ideal for scenarios like AI model training where multiple workers need access to the same dataset, or when you need to scale storage dynamically without downtime. Define shared file system using volumes in your compose specification: ``` volumes: file_system_data: driver: sharedFileSystem driver_opts: # efs configuration efsThroughputMode: provisioned efsPerformanceMode: generalPurpose efsProvisionedThroughputInMibps: 100 # filestore configuration filestoreCapacityGi: 1024 filestoreTier: "BASIC_HDD" # azure fileshare configuration fileshareQuotaGi: 1024 fileshareTier: Premium fileshareRedundancy: LRS # Nebius shared filesystem configuration nebiusFilesystemType: NETWORK_SSD nebiusFilesystemSizeGi: 1024 services: database: image: postgres volumes: - source: file_system_data target: /var/lib/redis/data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::EFS - source: file_system_data target: /var/lib/postgresql/data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::EFS gcp: clusterStorageType: GCP::FILESTORE azure: clusterStorageType: AZURE::FILE_SHARE nebius: clusterStorageType: NEBIUS::FILESYSTEM_NETWORK_SSD ``` **Configuration options:** - **driver**: Must be set to `sharedFileSystem` to enable Omnistrate shared file system - **driver_opts**: Customize the shared file system behavior - **efsThroughputMode**: EFS Throughput mode (provisioned, bursting) - **efsPerformanceMode**: EFS Performance mode (generalPurpose, maxIO) - **efsProvisionedThroughputInMibps**: EFS Provisioned throughput in MiB/s (when using provisioned mode) - **filestoreCapacityGi**: Filestore capacity in GiB - **filestoreTier**: Filestore tier - **filestoreMaxIopsPerTb**: Filestore throughput per Tb - **fileshareQuotaGi**: File Share quota in GiB - **fileshareTier**: File Share tier (Standard, Premium) - **fileshareRedundancy**: File Share redundancy (LRS, ZRS, GRS, GZRS) - **nebiusFilesystemType**: Nebius shared filesystem backend type (`NETWORK_SSD`, `NETWORK_HDD`, `WEKA`, `VAST`) - **nebiusFilesystemSizeGi**: Nebius shared filesystem size in GiB **Mount configuration:** - **source**: Name of the volume defined in the volumes section - **target**: Path where you want to mount the volume in the container - **type**: Must be set to `volume` - **x-omnistrate-storage**: Storage configuration with `clusterStorageType: AWS::EFS`, `clusterStorageType: GCP::FILESTORE`, `clusterStorageType: AZURE::FILE_SHARE`, or Nebius shared filesystem types such as `NEBIUS::FILESYSTEM_NETWORK_SSD`. For Nebius, the mount-level `clusterStorageType` must match the shared-volume backend selected in `driver_opts`: - `NETWORK_SSD` -> `NEBIUS::FILESYSTEM_NETWORK_SSD` - `NETWORK_HDD` -> `NEBIUS::FILESYSTEM_NETWORK_HDD` - `WEKA` -> `NEBIUS::FILESYSTEM_WEKA` - `VAST` -> `NEBIUS::FILESYSTEM_VAST` Omnistrate rejects Nebius shared file systems for `OMNISTRATE_MULTI_TENANCY` plans. ### volumes.driver: blob Blob storage allows you to define blob storage buckets and their mount paths across Resources. Omnistrate manages the lifecycle of storage volumes (creation and deletion with instance) and mounts them as local volumes. > **Note:** Blob storage is not available for the Hosted SaaS model. Define blob storage in the compose specification: ``` volumes: bucket_data: driver: blob services: app: image: myapp volumes: - source: bucket_data target: /mnt/blob-data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::S3 gcp: clusterStorageType: GCP::GCS azure: clusterStorageType: AZURE::BLOB_STORAGE ``` **Mount configuration:** - `bucket_data` is the name of the volume, you can opt for a name of your choice. - `driver: blob` is the driver type required to enable using Blob storage. - `x-omnistrate-storage` defines the type of storage for each cloud provider. This configuration will provision a new Blob storage (bucket/container) for each instance deployed. On the pod itself, it will appear just like any other volume. The underlying storage driver does not restrict or synchronize access, so your application itself needs to make sure it doesn't end up overriding data. ### x-omnistrate-capabilities `x-omnistrate-capabilities` allows you to add capabilities to your Resources. ``` services: web: image: nginx x-omnistrate-capabilities: httpReverseProxy: targetPort: 80 enableMultiZone: true enableEndpointPerReplica: true customDNS: targetPort: 80 autoscaling: maxReplicas: 1 minReplicas: 1 serverlessConfiguration: enableAutoStop: true minimumNodesInPool: 5 targetPort: 3306 networkType: INTERNAL # supported types are PUBLIC, and INTERNAL. If not specified, default is PUBLIC ``` Supported capabilities are custom DNS, custom sidecars, stable egress ip, process core dump, service account policies, backup configuration, autoscaling, serverless configuration, networkType. ### x-omnistrate-capabilities.customDNS The `customDNS` capability enables endpoint aliases for your Resources, allowing users to assign specific aliases to deployment or resource instance endpoints. This provides enhanced branding, customization, and control over infrastructure. ``` services: web: image: nginx x-omnistrate-capabilities: customDNS: targetPort: 80 ``` **Configuration fields:** - **targetPort**: The port number where your HTTP service is listening For application configuration that needs the customer-facing hostname, such as `ALLOWED_HOSTS`, CORS allow-lists, or callback URLs, use `$sys.network.customDNS` rather than node-level endpoints. ``` services: web: image: nginx environment: ALLOWED_HOSTS: "$sys.network.customDNS,localhost,127.0.0.1" x-omnistrate-capabilities: customDNS: targetPort: 80 ``` Behavior notes: - `$sys.network.customDNS` resolves to the configured custom DNS when present - if no custom DNS is configured, it falls back to `$sys.network.externalClusterEndpoint` - if custom DNS is added or changed after deployment, restart the instance so environment variables are refreshed ### x-omnistrate-capabilities.sidecars The `sidecars` capability allows you to enhance your SaaS product's functionality without changing its core service. Sidecars are add-on containers that run alongside your main application container, providing additional features or capabilities while maintaining isolation from the main application. ``` services: web: image: nginx x-omnistrate-capabilities: sidecars: tooling: imageNameWithTag: "busybox:stable" monitoring: imageNameWithTag: "prometheus/node-exporter:latest" securityContext: runAsUser: 10 runAsGroup: 99 runAsNonRoot: true capabilities: add: - SYS_RESOURCE resourceLimits: cpu: "250m" memory: "256Mi" command: - "/bin/node_exporter" args: - "--path.rootfs=/host" - "--collector.filesystem.ignored-mount-points" ``` **Configuration fields:** - **imageNameWithTag**: Container image with tag for the sidecar (required) - **securityContext**: Security context configuration for the sidecar - **runAsUser**: User ID to run the container - **runAsGroup**: Group ID to run the container - **runAsNonRoot**: Whether to run as non-root user - **capabilities**: Linux capabilities to add or drop - **resourceLimits**: Resource constraints for the sidecar - **cpu**: CPU limit (e.g., "250m" for 250 millicores) - **memory**: Memory limit (e.g., "256Mi" for 256 MiB) - **command**: Entry point command for the sidecar container - **args**: Arguments to pass to the command ### x-omnistrate-capabilities.enableStableEgressIP The `enableStableEgressIP` capability provides a stable egress IP address for outbound traffic from your Resources. This is useful when you need to whitelist your SaaS product's egress IP address with external services or APIs. ``` services: web: image: nginx x-omnistrate-capabilities: enableStableEgressIP: true ``` ### x-omnistrate-capabilities.processCoreDump The `processCoreDump` capability enables process core dump collection for debugging application crashes and analyzing process failures. ``` services: database: image: postgres x-omnistrate-capabilities: processCoreDump: /var/lib/data/cores/core.%e.%p.%t ``` **Configuration:** - Specify the path where core dumps should be stored - Use format specifiers for core dump file naming: - `%e`: Executable name - `%p`: Process ID - `%t`: Timestamp ### x-omnistrate-capabilities.serviceAccountPolicies The `serviceAccountPolicies` capability enables your application to securely access cloud-native services by configuring appropriate service account permissions. ``` services: app: image: myapp x-omnistrate-capabilities: serviceAccountPolicies: aws: - MSK_CONNECT - SECRETS_MANAGER - LAMBDA - SQS gcp: - WORKLOAD_IDENTITY_IAM_BINDING ``` **AWS Policies:** - **MSK_CONNECT**: Enables AWS MSK Connect access - **SECRETS_MANAGER**: Enables AWS Secrets Manager access - **LAMBDA**: Enables AWS Lambda access (including Serverless framework permissions) - **SQS**: Enables Amazon SQS access **GCP Policies:** - **WORKLOAD_IDENTITY_IAM_BINDING**: Binds resource workload identity to IAM service account, granting additional GCP permissions (Logs, Metrics, Secrets) ### x-omnistrate-capabilities.backupConfiguration The `backupConfiguration` capability enables automatic backup and point-in-time restore for your Resources, providing data protection and recovery capabilities. ``` services: database: image: postgres x-omnistrate-capabilities: backupConfiguration: backupRetentionInDays: 7 backupPeriodInHours: 2 snapshotBeforeDeletion: true # Optional: take final snapshot before deletion ``` **Configuration fields:** - **backupRetentionInDays** (required): Number of days to retain backups - **backupPeriodInHours** (required): Frequency of backup creation in hours - **snapshotBeforeDeletion** (optional): Controls whether a final manual snapshot is automatically created before resource deletion. Defaults to `false` if not specified. When enabled, a final snapshot will be taken before instance deletion unless explicitly skipped during the deletion operation. **Storage Volume Backup Control:** You can disable backups for specific storage volumes by adding `disableBackup: true` in the volume configuration: ``` services: database: image: postgres volumes: - source: temp_data target: /tmp/data type: volume x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP2 instanceStorageSizeGi: 50 disableBackup: true ``` > **Note:** Backup capability is only available for Tenant-Aware Resources and applies to AWS EBS, GCP Persistent Disk, and Azure Disk storage types. When backup retention is updated, changes only apply to new backups. ### x-omnistrate-capabilities.autoscaling For autoscaling based on custom application metrics instead of default CPU/memory metrics: ``` x-omnistrate-capabilities: autoscaling: scalingMetric: metricEndpoint: "http://localhost:9187/metrics" metricLabelName: "application_name" metricLabelValue: "psql" metricName: "pg_stat_activity_count" ``` `scalingMetric` expects a Prometheus endpoint that your application exposes. It does not scrape arbitrary authenticated HTTP APIs directly. If your scaling signal comes from a queue API, a protected service, or requires custom logic, use [Custom Autoscaling](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md) instead. ### x-omnistrate-capabilities.serverlessConfiguration For complex serverless scenarios requiring custom metrics, session state preservation, or proxy configuration control, you can use advanced serverless mode: ``` x-omnistrate-capabilities: autoscaling: maxReplicas: 5 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 serverlessConfiguration: targetPort: 3306 enableAutoStop: true minimumNodesInPool: 5 ``` Behavior notes: - basic serverless scales the entire deployment down to zero and wakes the entire deployment back up, even though the configuration is attached to a single Resource - in most cases, configure `serverlessConfiguration` on one representative Resource and port that reflects overall deployment traffic - custom DNS aliases do not currently trigger basic serverless wake-up; use the generated Omnistrate endpoint to resume a scaled-to-zero deployment ### x-omnistrate-load-balancer `x-omnistrate-load-balancer` allows you to configure a load balancer for your Resources in the compose specification file adding L7-L4 capabilities. Adding L7 load balancing capabilities to your Resources allows to route traffic based on the URL path, while L4 routes traffic based on the port number. ``` version: '3.9' x-omnistrate-load-balancer: https: - name: PGAdmin description: L7 Load Balancer for PGAdmin - New paths: - associatedResourceKey: admin path: / backendPort: 80 tcp: - name: Writer description: L4 Load Balancer for Writer ports: - associatedResourceKeys: - writer ingressPort: 5432 backendPort: 5432 services: admin: image: dpage/pgadmin4 writer: image: postgres reader: image: postgres ``` ### x-omnistrate-job-config `x-omnistrate-job-config` allows you to configure a Resource as a job that runs one time during create, update, or modification operations of a deployment. When a Resource is configured with job settings, it will execute once during the deployment lifecycle and complete its task. This is useful for initialization scripts, data migrations, setup tasks, or any one-time operations. ``` services: hello-world: build: context: ./jobs/hello-world dockerfile: Dockerfile environment: RAY_ADDRESS: "ray://{{ $var.rayClusterAddress }}:10001" SCRIPT_PATH: "submit_job.py" deploy: resources: reservations: cpus: "0.1" memory: 256M limits: cpus: "0.5" memory: 1G privileged: true platform: linux/amd64 volumes: - source: ./jobs/hello-world target: /app type: bind - source: ./tmp target: /tmp type: bind x-omnistrate-compute: replicaCountApiParam: numReplicas x-omnistrate-job-config: backoffLimit: 0 activeDeadlineSeconds: 3600 ``` Job-specific configurations: - **backoffLimit**: Number of retries before considering the job as failed (set to 0 to disable retries) - **activeDeadlineSeconds**: Maximum duration (in seconds) the job is allowed to run before termination > **Note:** Jobs are designed to run to completion and then terminate. They are not meant for long-running services. ### x-omnistrate-image-registry-attribute `x-omnistrate-image-registry-attributes` allows you to configure image registries. ``` x-omnistrate-image-registry-attributes: docker.io: auth: username: username password: password ``` > **Note:** For private image registries, username and password are required. Images will only be downloaded in your account (Hosted SaaS) or in your customers' account (BYOC). ### x-omnistrate-integrations (deprecated) `x-omnistrate-integrations` allows you to add integrations to your service. This tag is deprecated, use `x-customer-integrations`. ### x-customer-integrations `x-customer-integrations` allows you to configure customer-specific integrations such as licensing, logs, and metrics. ``` x-customer-integrations: logs: metrics: ``` ### x-customer-integrations.licensing The licensing protection system ensures that only authorized subscribed users can access and use your software. This is particularly important for BYOA/BYOC deployments and on-premises installations. ``` x-customer-integrations: licensing: licenseExpirationInDays: 7 productPlanUniqueIdentifier: 'PRODUCT-SAMPLE-SKU-UNIQUE-VALUE' ``` Configuration fields: - **licensing**: Configure licensing integration - **licenseExpirationInDays**: Number of days before license expires (default: 7 days) - **productPlanUniqueIdentifier**: Unique identifier for the product plan SKU (optional, defaults to product tier ID) ### x-customer-integrations.logs The following configuration enables collecting and exposing logs using an Omnistrate managed Prometheus server. ``` x-customer-integrations: logs: metrics: ``` You can choose to use a different provider to publish the logs to. When using BYOC (Bring Your Own Cloud) `provider: native` integrates with the cloud provider's native observability platform (CloudWatch for AWS, Operations Suite for GCP, Application Insights for Azure). End customers can visualize the logs on their own account. ``` x-customer-integrations: logs: provider: native metrics: provider: native ``` ### x-customer-integrations.metrics `x-customer-integrations.metrics` configures customer-facing metrics collection and dashboards. At a minimum: ``` x-customer-integrations: metrics: ``` For BYOC, use cloud-native metrics: ``` x-customer-integrations: metrics: provider: native ``` For Omnistrate native advanced metrics and dashboards: ``` x-customer-integrations: metrics: scrapeTargets: - name: postgres podSelector: app.kubernetes.io/name: postgresql app.kubernetes.io/instance: postgres endpoint: portNumber: 9187 path: /metrics scheme: http additionalMetrics: postgres: metrics: pg_stat_activity_count: "Active Sessions": aggregationFunction: sum labelFilters: datname: postgres state: active pg_locks_count: "Locks by Mode": aggregationFunction: sum groupByLabels: - mode labelFilters: datname: postgres dashboards: exporter: title: "Postgres Exporter" sections: - title: "Sessions and Locks" panels: - title: "Active Sessions" type: timeseries unit: short targets: - ref: "Active Sessions" - title: "Locks by Mode" type: timeseries unit: short targets: - ref: "Locks by Mode" ``` Field semantics: - `scrapeTargets`: where to scrape metrics from (`podSelector`/`serviceSelector` + endpoint). - `additionalMetrics..metrics`: models raw Prometheus metrics into stable series names. - `aggregationFunction`: `sum`, `min`, `max`, `avg`. - `groupByLabels`: keep selected dimensions in output. - `labelFilters`: include only matching label values. - `additionalMetrics..dashboards`: custom dashboard layout. - `sections[].panels[].targets[].ref` should reference series names defined under `metrics`. - Panel types: `timeseries`, `stat`, `gauge`, `table`. - `targets[].expr` is optional raw PromQL when needed. Dashboard property details: | Path | Type | Required | Description | | ------------------------------------ | ------------- | -------------------- | -------------------------------------- | | `dashboards..title` | string | No | Dashboard title override | | `dashboards..sections` | array | Yes | Ordered dashboard sections | | `sections[].title` | string | Yes | Section title | | `sections[].description` | string | No | Optional section context | | `sections[].panels` | array | Yes | Section panels | | `panels[].title` | string | Yes | Panel title | | `panels[].type` | enum | Yes | `timeseries`, `stat`, `gauge`, `table` | | `panels[].description` | string | No | Optional panel details | | `panels[].unit` | string | No | Unit format | | `panels[].min` / `panels[].max` | number | No | Value bounds | | `panels[].targets` | array | Yes | Series/query targets | | `targets[].ref` | string | Recommended | Reference to modeled series name | | `targets[].expr` | string | Optional alternative | Raw PromQL expression | | `targets[].legend` | string | No | Legend text override | | `targets[].axis` | enum | No | Axis side: `left` or `right` | | `panels[].thresholds` | array | No | Threshold list | | `thresholds[].color` | string | Yes (if used) | Threshold color | | `thresholds[].value` | number | Yes (if used) | Threshold boundary | | `panels[].tags` | array[string] | No | Panel tags | For copy/paste panel templates (health stat, saturation trend, distribution table), see [Build Guide / Integrations](https://docs.omnistrate.com/build-guides/integrations/#custom-metrics-and-dashboards-advanced). For legacy-to-advanced migration steps, see [Build Guide / Integrations](https://docs.omnistrate.com/build-guides/integrations/#migration-guide-legacy-custom-metrics-to-advanced-metrics-dashboards). ### x-internal-integrations `x-internal-integrations` allows you to configure internal integrations that are not exposed to customers. ``` x-internal-integrations: logs: metrics: ``` This extension allows you to define metrics and other integrations that are only visible internally and not exposed to your customers. #### Internal Observability Options **Omnistrate Native:** - **logs**: Enable real-time logging for your support team to manage the customer fleet - **metrics**: Enable real-time infrastructure metrics dashboard **OpenTelemetry Providers:** You can ship metrics and logs to third-party providers like NewRelic, Signoz, or Datadog: ``` x-internal-integrations: metrics: provider: newRelic # or signoz, datadog endpoint: https://otlp.nr-data.net secretLocators: aws: arn:aws:secretsmanager:us-west-2:xxxxxxxxxxx:secret:mySecret123456-abc123 gcp: projects/xxxxxxxxxxx/secrets/mySecret123456-abc123/versions/latest azure: KeyVaultName/SecretName serviceComponentsConfiguration: postgres: prometheusEndpoint: "http://localhost:9187/metrics" logs: provider: newRelic endpoint: https://otlp.nr-data.net secretLocators: aws: arn:aws:secretsmanager:us-west-2:xxxxxxxxxxx:secret:mySecret123456-abc123 ``` **Cloud Native:** Enable integration with cloud provider's native observability platform: ``` x-internal-integrations: logs: provider: native metrics: provider: native ``` #### Custom Metrics Configuration For additional application metrics, you can specify custom metrics with aggregation functions and label filters. You can then reference those modeled series in `dashboards`: ``` x-customer-integrations: metrics: additionalMetrics: postgres: prometheusEndpoint: "http://localhost:9187/metrics" metrics: pg_stat_activity_count: average_active_activity: aggregationFunction: avg labelFilters: state: active max_postgres_activity: aggregationFunction: max labelFilters: datname: postgres dashboards: exporter: title: Postgres Exporter sections: - title: Activity panels: - title: Active Sessions type: timeseries targets: - ref: average_active_activity ``` **Supported aggregation functions:** sum, avg, max, min ### x-internal-integrations.multiTenantGpu GPU slicing enables multiple workloads to efficiently share GPU resources, maximizing utilization and reducing costs. Omnistrate supports both NVIDIA time-slicing and Multi-Instance GPU (MIG) technologies. Enable GPU slicing using the `x-internal-integrations` extension: ``` x-internal-integrations: multiTenantGpu: instanceType: g4dn.xlarge timeSlicingReplicas: 2 migProfile: 1g.5gb # optional: for MIG-capable GPUs ``` **Configuration parameters:** - **instanceType**: GPU-enabled EC2 instance type (g4dn.xlarge, p3.2xlarge, p4d.24xlarge, etc.) - **timeSlicingReplicas**: Number of virtual GPU replicas (2, 4, 8, etc.) - **migProfile**: MIG profile for A100/H100 GPUs (optional) ## Variable Interpolation and System Parameters ### API Parameters Access API parameters defined in `x-omnistrate-api-params` using `$var.`: ``` services: database: image: postgres environment: - POSTGRES_PASSWORD=$var.postgresqlPassword - POSTGRES_DATABASE=$var.postgresqlDatabase - POSTGRES_USERNAME=$var.postgresqlUsername x-omnistrate-api-params: - key: postgresqlPassword type: String required: true - key: postgresqlDatabase type: String required: true - key: postgresqlUsername type: String required: true ``` ### Omnistrate System Parameters Omnistrate provides system parameters that are replaced with actual values at runtime using the format `$sys.`. These parameters provide contextual information about the deployment environment. > **Note:** All variables are substituted with their original data type. If you want to use a variable as a string, wrap it in escaped quotes: `\"$sys.deploymentCell.region\"`. ## Fragments You can use built-in YAML features to make your Compose file neater and more efficient. ### Anchors and Aliases Anchors are created using the `&` sign, and aliases use the `*` sign: ``` x-common-variables: &common-variables POSTGRES_DB: myapp POSTGRES_USER: postgres services: db: image: postgres environment: <<: *common-variables POSTGRES_PASSWORD: secret backup: image: postgres environment: *common-variables ``` # Compose Troubleshooting For the shared debugging workflow, start with [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md). It explains Debug Events, instance debug, and the relationship between workflow state and live resource state. ## Compose Troubleshooting Checklist Use `omnistrate-ctl instance debug ` to open the Compose resource and work through the following: 1. Review application logs for container startup, health-check, or dependency errors. 1. Confirm deployment API parameters resolved to the expected runtime values. 1. Confirm deployment output parameters contain the values other resources or customers expect. 1. Check workflow events to identify whether the failure happened during bootstrap, storage, network, compute, deployment, or monitoring. 1. Use the Metrics tab or `omnistrate-ctl instance dashboard ` to verify dashboard access when metrics are enabled. The shared [Compose Troubleshooting Checklist](https://docs.omnistrate.com/operate-guides/troubleshooting/#compose-troubleshooting-checklist) contains the same operational checks alongside the Helm, Operator, and Terraform checklists. ## Publishing Changes If you change the Compose specification, parameter mappings, source artifacts, or rendered configuration, publish a new Plan version and trigger a fresh workflow. Use [Workflows](https://docs.omnistrate.com/operate-guides/workflows/index.md) to distinguish a retry of the same released artifacts from a new version. # Dedicated Tenancy Deployment Guide This guide shows how to configure Dedicated Tenancy for Hosted SaaS deployments. ## Hosted SaaS For Compose-based SaaS Products, set `OMNISTRATE_DEDICATED_TENANCY` and configure the provider account with `hostedDeployment`: ``` x-omnistrate-service-plan: name: 'Dedicated-Tenancy SaaS' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' deployment: hostedDeployment: awsAccountId: '' awsBootstrapRoleAccountArn: 'arn:aws:iam:::role/omnistrate-bootstrap-role' ``` For complete isolation, you can also configure one deployment per cell with `maximumDeploymentsPerCell: 1`. See [Custom Deployment Cell Placement](https://docs.omnistrate.com/infra-guides/custom-deployment-cell-placement/index.md) and [Customer Networks](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md). ## What's next - [Dedicated Tenancy Overview](https://docs.omnistrate.com/build-guides/dedicated-tenancy-overview/index.md) - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/#dedicated-tenancy) - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) - [Customer Networks](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md) # Dedicated Tenancy Dedicated Tenancy gives each customer deployment its own isolated infrastructure stack (VMs). This model provides strong infrastructure, network, storage, and performance isolation while Omnistrate manages provisioning, scaling, and lifecycle operations. ## When to use Dedicated Tenancy This model is suitable for: - Enterprise and regulated customers. - Workloads with strict security or compliance requirements. - Applications that require predictable performance isolation. - Customers that need dedicated compute, storage, or network resources. For the complete isolation and placement guidance, see [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/#dedicated-tenancy). For configuration and rollout steps, see the [Dedicated Tenancy Deployment Guide](https://docs.omnistrate.com/build-guides/dedicated-tenancy-deployment/index.md). ## Related Guides - [Tenancy Types](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) - [Deployment Models](https://docs.omnistrate.com/build-guides/deployment-models/index.md) - [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) - [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) # Dependencies ## What are Resource dependencies Dependencies allow you to specify a DAG and how different Resources depend on each other. The dependency structure is in turn used to determine the order across Resources during several operations, such as provisioning, patching, and scaling. For example, if a parent Resource (say Kafka) depends on a child Resource (say ZooKeeper), Omnistrate provisions ZooKeeper before Kafka so that ZooKeeper is available when Kafka starts. Note - Dependencies are different from resource linking. To learn more about linking Resources, please see [this page](https://docs.omnistrate.com/build-guides/resource-linking/index.md). For comparison between the two, see [here](https://docs.omnistrate.com/build-guides/resource-linking/#resource-linking-vs-resource-dependency) - Note that some of the resources (Resources) can be internal in your DAG. To learn more about internal vs external resources, please see [this page](https://docs.omnistrate.com/build-guides/resource/#internal-vs-external-tenant-aware) - Note that some of the resources can be passive resources in your DAG. To learn more about passive vs active resources, please see [this page](https://docs.omnistrate.com/build-guides/resource/#passive-vs-active) ## Configuring dependencies You can define dependencies using `depends_on` tag in your compose specification. Here is a more concrete example: ``` services: PGAdmin: image: omnistrate/pgadmin4:7.5 x-omnistrate-mode-internal: true Writer: image: 'bitnami/postgresql:latest' x-omnistrate-mode-internal: true Reader: image: 'bitnami/postgresql:latest' x-omnistrate-mode-internal: true Cluster: image: omnistrate/noop depends_on: - Writer - Reader - PGAdmin ``` Alternatively, you can configure it using our API/CTL/UI. In the above example, Cluster resource depends on 3 resources: Writer, Reader and PGAdmin. The relationship between them will look like the following: Note Note that Writer, Reader and PGAdmin are all internal resources and so Omnistrate will NOT generate API/CLI/UI for your customers to directly provision or perform other management operations. Instead the lifecycle of these resources will be managed through Cluster parent resource, which is external. Note Also, note that Cluster is a passive resource and so creating an instance of just Cluster resource will NOT provision any infrastructure. However, Cluster resource depends on the other 3 resources and, so when create is invoked on the Cluster resource, it will create an instance of other 3 resources which are active resources with their associated infrastructure. ### Configuring dependency parameters As part of defining dependencies, you may also want to map dependency parameters. It allows you to map API parameters to configure dependent components to the parent component through dependency mapping. As an example: ``` services: Cluster: image: omnistrate/noop x-omnistrate-api-params: - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: true export: true defaultValue: t4g.small parameterDependencyMap: Writer: writerInstanceType Reader: readerInstanceType PGAdmin: instanceType ``` In the above example, the `instanceType` param of Cluster resource is mapped to `writerInstanceType` of Writer resource, `readerInstanceType` of Reader resource, and `instanceType` of PGAdmin resource. Note In case, you have two Resources, A and B, which rely on a third Resource, C. If they both define the parameterDependencyMap to Resource C, then depending on whether the user is creating an instance of A or of B, C will use the corresponding value. ### Dependency maps are not transitive `parameterDependencyMap` only applies from the parent resource to the direct child resource named in the map. It does not forward values through the rest of the dependency graph automatically. For example, if: - `App` depends on `Infra` - `App` depends on `SecureInfra` - `SecureInfra` also depends on `Infra` then any inputs required by `Infra` must still be mapped directly from `App` to `Infra`. Mapping a parameter from `App` to `SecureInfra` does not automatically populate `Infra`. ``` services: - name: Infra internal: true apiParameters: - key: access_type name: Access Type type: String required: false modifiable: true export: false defaultValue: public - name: SecureInfra internal: true dependsOn: - Infra apiParameters: - key: tenant_name name: Tenant Name type: String required: false modifiable: true export: false defaultValue: tenant-a - name: App dependsOn: - Infra - SecureInfra apiParameters: - key: accessType name: Access Type type: String required: false modifiable: true export: true defaultValue: public parameterDependencyMap: Infra: access_type - key: tenantName name: Tenant Name type: String required: false modifiable: true export: true defaultValue: tenant-a parameterDependencyMap: SecureInfra: tenant_name ``` If `SecureInfra` also needs outputs from `Infra`, model that with `dependsOn` and `{{ $Infra.out. }}` references. Input parameter mapping and output references solve different problems. # Deployment Cells: A Cellular Architecture for Multi-Cloud Service Delivery ## What is a Deployment Cell? Deployment cells are the foundational building blocks of Omnistrate's cellular architecture, designed to help SaaS Providers deliver and manage services across multiple clouds and customer environments. Think of deployment cells as isolated, self-contained units of compute infrastructure where your services run. Deployment cells leverage Kubernetes allowing you to leverage existing operational tools and practices while providing a robust framework for scaling, isolation, and operational efficiency. A deployment cell maps to a Kubernetes cluster plus the supporting network, system add-ons, and Omnistrate agents in a specific account and region. It is the data-plane location where your instances actually run. A deployment cell is **not** the Omnistrate control plane, SaaS portal, or API layer. ## Why Cellular Architecture? In traditional monolithic infrastructure, a single failure can cascade across your entire service delivery platform. Cellular architecture compartmentalizes your infrastructure into independent cells, providing: - **Blast Radius Containment**: Issues in one cell don't affect others - **Scalable Growth**: Add new cells as you grow without impacting existing deployments - **Geographic Distribution**: Deploy cells close to your customers for better performance - **Multi-Cloud Flexibility**: Mix and match cloud providers based on customer needs ## Key Concepts ### Cell Types and Use Cases #### Dedicated Customer Cells - Single-tenant infrastructure for enterprise customers - Complete isolation for compliance and security requirements - Custom configurations and deployment constraints - Ideal for: Regulated industries, high-security workloads, custom SLAs #### Shared Multi-Tenant Cells - Cost-effective infrastructure shared across multiple customers - Resource optimization through workload consolidation - Standardized configurations - Ideal for: SaaS Products, development environments, cost-sensitive workloads #### Adopted Cells - Bring-your-own-infrastructure model - Leverage existing customer Kubernetes clusters - Maintain customer's existing security and compliance postures - Ideal for: Enterprise customers with existing infrastructure investments ## Deployment Cell Lifecycle How a cell is created and reused depends on the deployment model: - **Hosted**: Omnistrate provisions the deployment cell in one of your cloud accounts and can place multiple instances there based on your tenancy and placement rules. - **BYOC**: Omnistrate provisions the deployment cell in the customer account. The first deployment in an account and region usually creates the cell, and later deployments in that account and region typically reuse it. - **Adopted**: The deployment cell is an existing Kubernetes cluster that you register with Omnistrate. Because cells are cluster-scoped infrastructure, not every operational dependency should be coupled to an individual instance lifecycle. ### Cell-wide prerequisites belong at the cell layer Install components that should exist once per cluster, such as shared Operators, Prometheus stacks, ExternalDNS, CSI drivers, ingress controllers, or policy controllers, at the deployment-cell layer. Use [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) for these prerequisites instead of treating them as manual post-install steps on each instance. This keeps cluster-wide dependencies upgradeable independently from any single tenant instance. ### Geographic Distribution Deployment cells naturally support geographic distribution: - Deploy cells in regions close to your customers - Maintain data residency compliance - Optimize for latency-sensitive applications - Enable disaster recovery through geographic redundancy ### Scaling Patterns #### Horizontal Scaling - Add new cells as customer base grows - No need to resize existing infrastructure - Zero-downtime expansion #### Customer-Driven Scaling - Start customers on shared cells - Graduate to dedicated cells as they grow ## Operational Benefits ### Isolation and Security - **Network Isolation**: Each cell has its own network boundaries - **Resource Isolation**: CPU, memory, and storage are cell-specific - **Failure Isolation**: Problems don't cascade across cells - **Security Isolation**: Compromised cells don't affect others ### Maintenance and Updates - **Rolling Updates**: Update cells independently - **Canary Deployments**: Test changes on specific cells first - **Maintenance Windows**: Schedule per-cell maintenance - **Version Management**: Different cells can run different versions temporarily ## Multi-Cloud Strategy Deployment cells enable true multi-cloud operations: ### Cloud Provider Flexibility - Run cells on AWS, Google Cloud, Azure, or OCI - Mix providers based on: - Customer preferences - Regional availability - Cost considerations - Feature requirements ### Hybrid Deployments - Combine cloud and on-premises cells - Support air-gapped environments ## Conclusion Deployment cells provide the architectural foundation for building scalable, resilient, and flexible service delivery platforms. By embracing cellular architecture, SaaS Providers can offer better isolation, performance, and reliability while maintaining operational efficiency across multiple clouds and customer environments. # Deployment Models ## What Deployment Models are supported by Omnistrate We support three Deployment Models to decide where your application runs and meet your distribution needs: - **Hosted SaaS**: Deployed in your SaaS provider account - **BYOC Anywhere**: Deployed in your customer's account - **Air-Gapped**: Deployed to isolated cloud environments and local hardware For a comprehensive overview of all deployment scenarios and use cases, visit our [Use Cases](https://docs.omnistrate.com/usecases/overview/index.md) section. For Hosted SaaS tenancy options, see the [Cellular Multi-Tenancy Overview](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-overview/index.md) and [Dedicated Tenancy Overview](https://docs.omnistrate.com/build-guides/dedicated-tenancy-overview/index.md). ## Hosted SaaS The Hosted SaaS model deploys your application in your SaaS provider account and provides a fully managed experience to your customers. Many SaaS Products like Slack, Stripe, GitHub are offered in this model. Here is a reference architecture: ### Customer Networks For enhanced security and network isolation while maintaining this deployment model, you can enable **Customer Networks**. This advanced feature allows your customers to define network partitioning on dedicated stacks, providing complete isolation and the ability to set up private network paths for customer connectivity. This gives you the benefits of dedicated infrastructure per customer while keeping services self-served through Hosted SaaS. To learn more about configuring customer networks, see [Customer Networks](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md). ## BYOC Anywhere Your customers require that data stays in their account due to security and migration cost. To support them, you have to host your application in their account as a fully-managed solution. The BYOC deployment model exactly addresses this use-case by seamlessly establishing the trust relationship between your account and your customers' accounts enabling the deployment and management of resources. We follow the industry standard secure techniques to reverse the connection to prevent any inbound connections to your customers' account, encrypted channel through TLS and OAuth to secure the connectivity between your customers account and your account. Many SaaS Products like Databricks, RedPanda BYOC are offered in this model. Note There are several terms for BYOC mode in the industry and they are all somewhat related. - Bring Your Own Cloud (BYOC) - Bring Your Own VPC Here is a reference architecture: To learn more, check the [BYOC Anywhere](https://docs.omnistrate.com/usecases/byoc/index.md) use case. ## Air-Gapped Air-Gapped installations are suitable for isolated environments, running on a company's local hardware. This model offers the greatest control, security, and independence but requires high initial costs and internal maintenance. In contrast, cloud installations run on a provider's remote servers, offering significant cost savings through a subscription model, enhanced scalability, flexibility, and faster deployment, though at the cost of less direct control and reliance on the internet and the provider. Here is a reference architecture: Learn more about air-gapped deployments in the [Air-Gapped Use Case](https://docs.omnistrate.com/usecases/air-gapped/index.md). For a step-by-step guide to building an air-gapped installer, see [Air-Gapped Installer Plan](https://docs.omnistrate.com/build-guides/air-gapped-helm-charts/index.md). # Expression evaluation ## Evaluating Expressions The `omctl instance evaluate` subcommand introduces a powerful new capability to the Omnistrate CLI, allowing you to **evaluate Omnistrate expressions directly within the context of an active deployment or instance.** This feature is designed to significantly accelerate your development and debugging workflows, especially when working with Terraform or Helm charts that rely heavily on Omnistrate expressions and variable references. ### Why do we need to Evaluate Expressions? Before this command, iterating on changes in your infrastructure code (like Terraform or Helm) that utilized Omnistrate expressions often involved a cycle of deploying, then inspecting logs or outputs to verify expression results. This process could be time-consuming and cumbersome. With `omctl instance evaluate`, you can now: - **Rapidly Prototype Expressions:** Quickly test and refine your Omnistrate expressions without needing a full deployment cycle. - **Debug Variable References:** Instantly check the resolved values of `$var` (instance-specific variables), `$sys` (system parameters), and other expression components in a live instance context. - **Validate Configuration:** Ensure that your expressions will resolve to the expected values before applying them in your deployment definitions. - **Understand Instance State:** Gain immediate insights into the runtime configuration and system parameters of any given instance. This command transforms the development experience, making it much faster and more efficient to work with Omnistrate's dynamic expression capabilities. ## Key Features to Evaluate Expressions - **Single Expression Support:** Evaluate individual expressions directly from the command line. - **Expression File Support:** Load and evaluate multiple expressions defined in a JSON file. - **Mutual Exclusivity:** Ensure clarity by using either the `--expression` or `--expression-file` flag, but not both. - **Full Output Format Support:** View results in `table`, `text`, or `json` formats, consistent with other `omctl` commands. ## How to Evaluate Expressions The `evaluate` command requires an `[instance-id]` and a `[resource-key]`. **Evaluate a single expression:** ``` omctl instance evaluate instance-ab1c2d3e terraform-infra --expression '{{ $var.username }} {{ $sys.id }}' ``` **Evaluate multiple expressions from a JSON file:** First, create a `expressions.json` file like this: ``` { "greeting": "Hello {{ $sys.tenant.name }}" } ``` Then, run the command: ``` omctl instance evaluate instance-fg4h5i6j helm-chart --expression-file expressions.json ``` **Specify output format:** ``` omctl instance evaluate instance-kl7m8n9o terraform-infra --expression 'test' --output json ``` ### Arguments and Flags - **`[instance-id]`** (Required): The unique identifier of the instance you want to evaluate expressions against. - **`[resource-key]`** (Required): The key of the resource within the instance where the expressions will be evaluated. - **`--expression`, `-e`**: (String) The expression string to evaluate. Use this for single expressions. - **`--expression-file`, `-f`**: (Path) Path to a JSON file containing a map of expressions to evaluate. Use this for multiple expressions. - **`--output`, `-o`**: (String) The desired output format: `table`, `text`, or `json`. (Inherited from global flags) ### Discovering available fields When you are not sure which fields are populated for a given resource, start by inspecting the top-level system objects directly: ``` omctl instance evaluate --expression '$sys.deploymentCell' -o json | jq omctl instance evaluate --expression '$sys.deployment' -o json | jq omctl instance evaluate --expression '$sys.compute' -o json | jq omctl instance evaluate --expression '$sys.network' -o json | jq ``` For Terraform or OpenTofu resources, prefer starting with `$sys.deploymentCell.*`. Those resources are evaluated before workload-specific node state is available, so node-scoped fields such as `$sys.compute.node.*` or `$sys.network.node.*` may be empty or unavailable there. `$sys.deploymentCell.kubernetesClusterID` is the cloud provider's cluster identifier. It can differ from the Omnistrate deployment cell ID. ### Output Formats The output format adapts based on the `--output` flag. For a single expression: ``` { "result": "evaluated_value_here" } ``` For expression maps: ``` { "result": { "greeting": "Hello Jane Doe" } } ``` ## Evaluate Expressions Examples Here are some practical examples of using `omctl instance evaluate`: **Evaluate system parameters:** ``` ➜ omctl instance evaluate instance-qr0s1t2u terraform-infra --expression '$sys.deploymentCell.oidcIssuerID' -o json | jq { "result": "eastus2.oic.prod-aks.azure.com/1a2b3c4d-e5f6-7g8h-9i0j-1k2l3m4n5o6p/7q8r9s0t-1u2v-3w4x-5y6z-7a8b9c0d1e2f/" } ``` **Evaluate expressions to check instance deployment parameters:** ``` ➜ omctl instance evaluate instance-vw3x4y5z helm-chart --expression '$var.gpu_enabled' -o json | jq { "result": true } ``` **Evaluate complex expressions combining variables and system parameters:** ``` ❯ omctl instance evaluate instance-j8k9l0m1 helm-chart --expression 'GPU Enabled: {{ $var.gpu_enabled }} on {{ $sys.deploymentCell.cloudProviderName }} for {{ $sys.tenant.name }}' -o json | jq { "result": "GPU Enabled: true on aws for Jane Doe" } ``` # Layered Chart Values ## What is a Layered Chart Values Omnistrate supports layered chart values for Helm chart configuration, providing conditional deployment customization across different environments, cloud providers, and deployment contexts. This feature replaces traditional static `chartValues` with a flexible, hierarchical system that applies configuration layers based on runtime conditions. Info Layered Chart Values and traditional `chartValues` are mutually exclusive. You must choose one approach per resource. ## Basic example for Layered Chart Values ``` layeredChartValues: - # Base layer - applies to all deployments values: global: environment: "production" monitoring: enabled: true - # Conditional layer - applies only to AWS scope: "{{ $sys.deploymentCell.cloudProviderName }}": "aws" values: aws: storageClass: "gp3" loadBalancer: type: "nlb" - # External configuration layer scope: "{{ $sys.deploymentCell.region }}": "us-west-2" valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/helm-configs.git" reference: "refs/heads/main" accessToken: "{{ $secret.GH_ACCESS_TOKEN }}" path: "regions/us-west-2/values.yaml" ``` ## How Layered Chart Values work Layers are processed in definition order. Each layer's scope conditions are evaluated against deployment context, and only matching layers are applied. Configuration values are deep merged, with later layers overriding earlier ones. **Template variables available:** - **System variables** (`$sys.*`): Platform parameters like cloud provider, region, network config. See [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md) - **API parameters** (`$var.*`): User-defined API parameters from your service specification - **Output variables** (`$out.*`): Infrastructure outputs from Terraform resources and other resource outputs - **Secrets** (`$secret.*`): Secure access to secrets managed by Omnistrate ## Layered Chart Values Types ### Inline values layer Defines configuration directly in the service specification: ``` - values: app: name: "my-service" replicas: 3 image: tag: "v1.2.3" ``` ### Git-referenced layer References external configuration files: ``` - valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/configs.git" reference: "refs/heads/main" accessToken: "{{ $secret.GITHUB_TOKEN }}" path: "services/my-service/values.yaml" ``` ### Scoped layer Applies only when conditions are met: ``` - scope: "{{ $sys.deploymentCell.cloudProviderName }}": "gcp" "{{ $sys.deploymentCell.region }}": "us-central1" values: gcp: workloadIdentity: enabled: true ``` ## Advanced examples of Layered Chart Values ### Multi-cloud deployment ``` layeredChartValues: # Base configuration for all clouds - values: app: name: "multi-cloud-service" replicas: 3 monitoring: enabled: true # AWS-specific configuration - scope: "{{ $sys.deploymentCell.cloudProviderName }}": "aws" values: serviceAccount: annotations: eks.amazonaws.com/role-arn: "{{ $out.iam.role.arn }}" storage: class: "gp3" provisioner: "ebs.csi.aws.com" # GCP-specific configuration - scope: "{{ $sys.deploymentCell.cloudProviderName }}": "gcp" values: serviceAccount: annotations: iam.gke.io/gcp-service-account: "{{ $out.iam.serviceAccount.email }}" storage: class: "ssd" provisioner: "pd.csi.storage.gke.io" # Azure-specific configuration - scope: "{{ $sys.deploymentCell.cloudProviderName }}": "azure" values: serviceAccount: annotations: azure.workload.identity/client-id: "{{ $out.iam.identity.clientId }}" storage: class: "managed-csi" provisioner: "disk.csi.azure.com" ``` ### Environment-based configuration ``` layeredChartValues: # Production defaults - values: app: replicas: 5 resources: cpu: "1000m" memory: "2Gi" monitoring: level: "standard" # Development overrides - scope: "{{ $var.environment }}": "development" values: app: replicas: 1 resources: cpu: "100m" memory: "256Mi" monitoring: level: "debug" # Staging configuration from Git - scope: "{{ $var.environment }}": "staging" valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/staging-configs.git" reference: "refs/heads/main" accessToken: "{{ $secret.GH_ACCESS_TOKEN }}" path: "common/staging-values.yaml" ``` ### Large customer-specific override blocks When the number of customer-specific Helm overrides is too large to model as individual API parameters, define a single `json` API parameter and merge it as a layered values block. ``` services: - name: MyApp apiParameters: - key: customChartValues name: Custom Chart Values description: Customer-specific Helm overrides for advanced cases type: json modifiable: true required: false export: true defaultValue: | metrics: serviceMonitor: enabled: true helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-charts chartRepoURL: https://charts.example.com layeredChartValues: - values: global: environment: production - name: customer-overrides values: $var.customChartValues ``` This pattern is useful when: - You need a flexible escape hatch for advanced customer overrides - The override surface is too large to register field-by-field as API parameters - You want a stable base layer with a single customer-provided override layer on top Note The `json` parameter should resolve to an object-shaped YAML or JSON document that matches the Helm values structure you want to merge. Use this for maps of values, not for arbitrary free-form text. ## Git integration for Layered Chart Values ### Repository authentication #### Public repositories For publicly accessible repositories, only specify the repository URL: ``` valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/public-helm-configs.git" reference: "refs/heads/main" path: "shared/base-values.yaml" ``` #### Private repositories For private repositories, provide an access token via Omnistrate secrets: ``` valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/private-helm-configs.git" reference: "refs/heads/main" accessToken: "{{ $secret.GITHUB_TOKEN }}" path: "secure/production-values.yaml" ``` ### Reference types Omnistrate supports standard Git reference formats for specifying which version of your configuration to use: #### Branch references ``` # Latest commit on main branch reference: "refs/heads/main" # Latest commit on develop branch reference: "refs/heads/develop" # Latest commit on feature branch reference: "refs/heads/feature/new-config" ``` #### Tag references ``` # Specific release tag reference: "refs/tags/v1.2.3" # Latest stable release reference: "refs/tags/stable" ``` #### Commit references For referencing specific commits, use the `commitSHA` field instead of `reference`: ``` valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/helm-configs.git" reference: "refs/heads/main" commitSHA: "a1b2c3d4e5f67890xxxxxxxxxxx34567890abcd" accessToken: "{{ $secret.GH_ACCESS_TOKEN }}" path: "values.yaml" ``` # Helm chart customization ## How to customize Helm Chart Values In the [Getting Started](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) guide, we designed a SaaS Product that deploys Redis Clusters using a Helm chart. However, in that example, we had all the pods deploy to any available VM in the Kubernetes cluster. In addition, these pods were collocated on the same VM, not an ideal scenario for High Availability. In this guide, we will show you how to customize the infrastructure for the Helm chart to deploy the Redis Master and Replica pods on separate VMs. We will also introduce a custom instance type for the Redis Master and Replica pods that your customers can specify through the Customer Portal we generated. Info A complete description of the Plan specification can be found on [Getting started / Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) ## Reference API Params from Helm Chart Values First, we need to update the Helm chart to support the new features. We will add two new parameters to the `values.yaml` file to allow customers to choose the instance type for their workload and the number of replicas they want to deploy. Info You can use system parameters to customize Helm Chart values. A detailed list of system parameters be found on [Build Guide / System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ``` ... apiParameters: - key: replicas description: Number of Replicas name: Replica Count type: Float64 modifiable: true required: false export: true defaultValue: "1" - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" ... ``` The `apiParameters` section in the specification defines the API parameters that you want your customers to provide as part of the provisioning APIs for your SaaS. For more information see: [API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md). Next, we will use the new API parameters to set the number of replicas and the instance type for the Redis Master and Replica pods in the `values.yaml` file. ``` ... compute: instanceTypes: - apiParam: instanceType cloudProvider: aws - apiParam: instanceType cloudProvider: gcp - apiParam: instanceType cloudProvider: azure replica: ... replicaCount: $var.replicas ... ``` Note For AWS EKS workloads that need EBS volumes encrypted with a customer-managed AWS KMS key, create a custom storage class and reference it from your Helm values. See [AWS Storage Classes](https://docs.omnistrate.com/infra-guides/storage-classes/#aws-storage-classes). ## Configure High Availability with Affinity Rules To ensure High Availability, you need to deploy pods across separate VMs. Omnistrate provides two approaches for managing Kubernetes affinity rules in Helm deployments: 1. **Automatic injection (default)** — Omnistrate automatically injects the required node affinity and pod anti-affinity rules into your Helm chart manifests at deploy time. 1. **Manual configuration** — You explicitly define affinity rules in your Helm chart values for full control. ### Automatic affinity injection (default) By default, Omnistrate automatically injects affinity rules into all applicable Kubernetes workload resources (Deployments, StatefulSets, DaemonSets, ReplicaSets, Jobs, CronJobs, and Pods) rendered by your Helm chart. This means you **do not need to manually specify affinity rules** in your chart values — Omnistrate handles it for you. The automatic injection performs the following: - **Node affinity**: Ensures pods are scheduled only on Omnistrate-managed nodes in the correct region, matching the designated resource and node pool version. The injected node selector expressions include: - `omnistrate.com/managed-by` — targets Omnistrate-managed nodes - `topology.kubernetes.io/region` — matches the deployment region - `omnistrate.com/resource` — matches the resource ID - `omnistrate.com/version` — matches the node pool version (`$sys.compute.node.version`) - **Pod labels**: Adds the `omnistrate.com/schedule-mode: exclusive` label to pod templates. - **Pod anti-affinity**: Ensures pods with the `exclusive` schedule mode are spread across different hosts for High Availability. Note Automatic injection is **enabled by default**. If your chart values already contain affinity rules, Omnistrate intelligently merges the injected rules with your existing configuration without creating duplicates. This behavior is controlled by the `chartAffinityControl` setting in the `helmChartConfiguration` section: ``` ... helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartAffinityControl: enableInjection: true # Enabled by default enableSharedHost: true # Enabled by default ... ``` | Property | Type | Default | Description | | ----------------------------------------------------------------------------------------------------------------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enableInjection` | boolean | `true` | Enables automatic injection of node affinity and pod anti-affinity rules into Helm chart manifests. | | `enableSharedHost` | boolean | `true` | Controls whether pods that use automatic affinity injection are allowed to share the same node. When set to `false`, Omnistrate enforces strict anti-affinity so | | | | | those pods are scheduled on separate hosts whenever cluster capacity allows, rather than being co-located for bin-packing. | | To **disable** automatic affinity injection (for example, if you want full manual control), set `enableInjection` to `false`: | | | | ``` ... chartAffinityControl: enableInjection: false ... ``` ### Manual affinity configuration (advanced) If you prefer to manage affinity rules entirely yourself, you can disable automatic injection and define the rules directly in your chart values. However, this is **rarely necessary** — automatic injection intelligently merges with any custom affinity rules you define in your chart values without overriding them. Only Omnistrate-specific rules that are not already present are appended, so your custom affinity logic is always preserved. Tip You do **not** need to disable automatic injection to use custom affinity rules. Omnistrate merges its rules with yours, so you can add custom affinity logic in your chart values and leave injection enabled. First, add a pod label for setting the anti-affinity between the master and replica pods. ``` ... master: podLabels: omnistrate.com/schedule-mode: exclusive replica: podLabels: omnistrate.com/schedule-mode: exclusive ... ``` Then, set the affinity rules for the Redis Master and Replica pods. ``` ... affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID - key: omnistrate.com/version operator: In values: - $sys.compute.node.version podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname ... ``` You might have noticed the use of system parameters (`$sys.deploymentCell.region`, `$sys.compute.node.instanceType`, `$sys.deployment.resourceID`, and `$sys.compute.node.version`) in the affinity rules. These variables are used to dynamically set the affinity rules based on the customer's deployment configuration. For more information, see [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ## Example specification with customized Helm Chart Values The following example uses the default automatic affinity injection, so no manual affinity rules are needed in the chart values: ``` name: Redis Server # Plan Name deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: Redis Cluster compute: instanceTypes: - apiParam: instanceType cloudProvider: aws - apiParam: instanceType cloudProvider: gcp - apiParam: instanceType cloudProvider: azure network: ports: - 6379 helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartValues: master: persistence: enabled: false service: type: LoadBalancer annotations: external-dns.alpha.kubernetes.io/hostname: $sys.network.externalClusterEndpoint resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi replica: persistence: enabled: false replicaCount: $var.replicas resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi apiParameters: - key: replicas description: Number of Replicas name: Replica Count type: Float64 modifiable: true required: false export: true defaultValue: "1" - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" ``` Note Compare this with the manual approach — you no longer need to specify `podLabels`, `nodeAffinity`, or `podAntiAffinity` in the chart values. Omnistrate injects the appropriate rules automatically at deploy time. Example with manual affinity rules (automatic injection disabled) If you need full manual control over affinity rules, disable automatic injection and define the rules explicitly: ``` name: Redis Server # Plan Name deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: Redis Cluster compute: instanceTypes: - apiParam: instanceType cloudProvider: aws - apiParam: instanceType cloudProvider: gcp - apiParam: instanceType cloudProvider: azure network: ports: - 6379 helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartAffinityControl: enableInjection: false chartValues: master: podLabels: omnistrate.com/schedule-mode: exclusive persistence: enabled: false service: type: LoadBalancer annotations: external-dns.alpha.kubernetes.io/hostname: $sys.network.externalClusterEndpoint resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID - key: omnistrate.com/version operator: In values: - $sys.compute.node.version podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname replica: podLabels: omnistrate.com/schedule-mode: exclusive persistence: enabled: false replicaCount: $var.replicas resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID - key: omnistrate.com/version operator: In values: - $sys.compute.node.version podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname apiParameters: - key: replicas description: Number of Replicas name: Replica Count type: Float64 modifiable: true required: false export: true defaultValue: "1" - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" ``` ### Apply changes to the service For this we will run the same command that was used to setup the service the first time. ``` omnistrate-ctl build -f spec.yaml --name 'RedisHelm' --release-as-preferred --spec-type ServicePlanSpec # Example output shown below ✓ Successfully built service Check the Plan result at: https://omnistrate.cloud/product-tier?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm Access your SaaS Product at: https://saasportal.instance-w6vidhd14.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm ``` ### Deploying a Redis Cluster through your dedicated Customer Portal Now, your customers can deploy Redis Clusters with the desired instance type and number of replicas through the dedicated customer portal. The portal will use the REST API parameters to customize the deployment based on the customer's requirements. **The final cluster status** **And the workload status in the Kubernetes Dashboard confirming the placement of pods on independent VMs** # Helm Charts Overview ## What is Helm? Helm is the package manager for Kubernetes, often described as the "apt/yum/homebrew for Kubernetes." It simplifies the deployment and management of applications on Kubernetes clusters by packaging complex applications into reusable, versioned charts. ### Key Concepts **Helm Charts** are packages of pre-configured Kubernetes resources that define, install, and upgrade even the most complex Kubernetes applications. Think of them as blueprints that contain: - **Templates**: Kubernetes manifest files with placeholders for dynamic values - **Values**: Configuration parameters that customize the deployment - **Dependencies**: Other charts that your application depends on - **Metadata**: Information about the chart version, description, and maintainers **Releases** are running instances of a chart deployed to a Kubernetes cluster. Each release has a unique name and can be upgraded, rolled back, or deleted independently. ## Why Use Helm Charts? ### Simplified Deployment Management Instead of managing dozens of individual Kubernetes YAML files, Helm charts package everything into a single, manageable unit. This reduces complexity and eliminates the risk of missing dependencies or configuration files. ### Parameterization and Reusability Helm charts use templating to make deployments configurable. The same chart can be deployed with different configurations for development, staging, and production environments, or customized for different customers. ### Version Control and Rollbacks Helm tracks deployment history, making it easy to upgrade applications or roll back to previous versions when issues arise. This provides confidence when deploying updates to production systems. ### Ecosystem and Community Helm has a rich ecosystem with thousands of pre-built charts available through repositories like [Artifact Hub](https://artifacthub.io/). Popular applications like PostgreSQL, Redis, Nginx, and Prometheus have well-maintained charts ready for use. ### Dependency Management Charts can declare dependencies on other charts, automatically handling complex multi-component applications. For example, a web application chart might depend on a database chart and a caching layer chart. ## Common Use Cases for Helm Charts ### Application Deployment Deploy complex applications with multiple components, services, and configurations in a single command. ### Environment Management Use the same chart with different values files to deploy identical applications across development, staging, and production environments. ### Multi-Tenancy Deploy multiple instances of the same application for different customers or teams, each with their own configuration and resources. ### CI/CD Integration Integrate Helm into continuous deployment pipelines to automate application updates and rollouts. ## Helm Integration with Omnistrate Omnistrate provides native support for Helm charts, making it incredibly easy to transform your existing Kubernetes applications into fully-managed SaaS offerings. Here's how Helm charts integrate seamlessly with the Omnistrate platform: ### **Zero-Modification Deployment** Bring your existing Helm charts to Omnistrate without any modifications. Whether you're using charts from public repositories like Bitnami or custom charts developed in-house, Omnistrate can deploy them directly. ### **Automated Infrastructure Management** When you deploy a Helm chart through Omnistrate, the platform automatically handles: - **Cloud Infrastructure**: VPCs, subnets, security groups, and networking - **Kubernetes Clusters**: Fully managed, production-ready clusters - **Load Balancing**: Network load balancers with Nginx ingress controllers - **DNS Management**: Route53 hosted zones for custom endpoints - **TLS Certificates**: Automated ACME certificate provisioning and rotation - **Monitoring**: Kubernetes dashboard and observability tools - **IAM/RBAC**: Proper service accounts and role bindings ### **Multi-Cloud and Multi-Tenant Ready** Omnistrate deploys your Helm charts across AWS, Google Cloud, and Azure, handling the cloud-specific configurations automatically. Each customer deployment is isolated with proper multi-tenancy controls. ### **Dynamic Configuration** Leverage Omnistrate's system parameters and API parameters to make your Helm charts dynamically configurable: ``` # Example: Dynamic instance types and replica counts replica: replicaCount: $var.replicas compute: instanceTypes: - apiParam: instanceType cloudProvider: aws ``` ### **Built-in SaaS Features** Omnistrate automatically adds enterprise SaaS capabilities to your Helm chart deployments: - **Customer Portals**: Self-service deployment interfaces - **REST APIs**: Programmatic access for customer integrations - **Billing Integration**: Usage metering and billing provider connections - **Backup and Recovery**: Automated backup strategies - **Monitoring and Alerting**: Customer-facing observability - **Upgrade Management**: Controlled rollouts and rollbacks ### **Simplified Operations** Instead of managing Kubernetes clusters, Helm releases, and infrastructure across multiple clouds, Omnistrate provides a unified control plane where you can: - Deploy and manage Helm chart-based services - Monitor customer deployments across all clouds - Handle upgrades and maintenance operations - Manage customer access and billing ## Getting Started with Helm Charts on Omnistrate The process of deploying Helm charts on Omnistrate is straightforward: 1. **Define Your Service**: Create a service specification that references your Helm chart 1. **Configure Parameters**: Set up API parameters for customer customization 1. **Deploy and Test**: Use the development environment to validate your deployment 1. **Go Live**: Release your Helm chart-based SaaS product to customers Here's a simple example of a Redis service using a Helm chart: ``` name: Redis Server services: - name: Redis Cluster helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartValues: master: persistence: enabled: false service: type: LoadBalancer replica: replicaCount: 1 persistence: enabled: false ``` This specification automatically creates: - A customer portal for Redis deployments - REST APIs for programmatic access - Multi-cloud deployment capabilities - Monitoring and management interfaces ## Error Handling and Stability Omnistrate's Helm integration includes built-in error handling and stability features that ensure reliable deployments across your fleet. ### Pre-install validation Before executing a Helm install, Omnistrate validates the chart configuration, template rendering, and resource requirements. Errors such as invalid chart values, missing dependencies, or template rendering failures are reported immediately — before any Kubernetes resources are created. This fail-fast approach prevents partially deployed releases that are difficult to clean up. ### First-install failure handling For the initial installation of a Helm release, if the first install fails, Omnistrate immediately cleans up the release and does not leave it in a pending or failed state. This ensures a clean slate for retry and avoids issues with Helm releases stuck in an inconsistent state. ### Improved error reporting When a Helm operation fails, Omnistrate captures and surfaces detailed error information including: - **Helm client logs** — the full output from the Helm install or upgrade operation - **Chart rendering errors** — template syntax issues or missing values - **Kubernetes resource errors** — pod scheduling failures, image pull errors, and resource quota issues - **Timeout details** — which resources failed to become ready within the configured timeout You can access these details through [Debug Events](https://docs.omnistrate.com/operate-guides/troubleshooting/#debug-events) in the instance details page, or by using the `omctl instance debug` command for direct access to Helm client logs and chart values. ### Agent crash recovery If the Omnistrate dataplane agent crashes during a Helm operation, the system automatically recovers the deployment state on restart. The agent reconciles the expected state with the actual Kubernetes state, resuming or retrying operations as needed without requiring manual intervention. You can control this behavior using the [`disableReconciliation`](https://docs.omnistrate.com/build-guides/helm-charts-runtime-configuration/index.md) flag in the runtime configuration. ### Best practices for stable deployments | Practice | Configuration | | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | Set appropriate timeouts | `timeoutNanos` in [runtime configuration](https://docs.omnistrate.com/build-guides/helm-charts-runtime-configuration/index.md) | | Wait for resources to be ready | `wait: true` and `waitForJobs: true` | | Disable problematic hooks | `disableHooks: true` for debugging | | Use fail-fast mode (debugging only) | `disableReconciliation: true` to disable crash recovery and automatic retries; not recommended for normal production use | | Manage CRDs separately | `skipCRDs: true` if CRDs are managed by an Operator | ## Next Steps Ready to start building with Helm charts on Omnistrate? Here are the recommended next steps: - **[Build from Helm](https://docs.omnistrate.com/getting-started/build-from-helm/index.md)**: Step-by-step guide to deploy your first Helm chart-based service - **[Helm Chart Customization](https://docs.omnistrate.com/build-guides/helm-charts-customize/index.md)**: Learn how to customize charts for high availability and customer requirements - **[System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md)**: Understand how to use dynamic parameters in your Helm charts - **[Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md)**: Advanced configuration techniques for complex deployments ## Benefits Summary By using Helm charts with Omnistrate, you get: ✅ **Rapid SaaS Development**: Transform existing applications into SaaS offerings in hours, not months ✅ **Enterprise-Grade Infrastructure**: Production-ready deployments with security, monitoring, and compliance built-in ✅ **Multi-Cloud Flexibility**: Deploy across AWS, GCP, and Azure without cloud-specific expertise ✅ **Operational Simplicity**: Unified management interface for all customer deployments ✅ **Customer Self-Service**: Automated portals and APIs for customer onboarding and management ✅ **Scalable Architecture**: Handle growth from single customers to thousands of deployments The combination of Helm's packaging power and Omnistrate's SaaS platform capabilities provides the fastest path from Kubernetes application to production SaaS business. # Helm Charts Runtime Configuration ## Overview Helm Charts Runtime Configuration allows you to control how Helm packages are deployed and managed within Omnistrate. These configuration options correspond to the runtime behavior of Helm operations, providing fine-grained control over deployment strategies, timeouts, hooks, and upgrade behaviors. The runtime configuration is specified using the `helmRuntimeConfiguration` struct on a service specification and can be applied during service plan creation or updates to customize how your Helm charts are processed. Info Runtime configuration parameters are optional and have sensible defaults. Only specify parameters when you need to override the default behavior for your specific use case. ## Configuration Parameters ### DisableHooks **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--no-hooks` Controls whether Helm hooks are executed during deployment operations. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: disableHooks: true ``` **When to use:** - Skip pre/post-install, pre/post-upgrade, or pre/post-delete hooks - Faster deployments when hooks are not required - Troubleshooting deployment issues caused by failing hooks **Example use case:** If your Helm chart includes database migration hooks that are causing deployment failures, you can disable hooks temporarily to deploy the main application first. ### Wait **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--wait` Determines whether Helm waits for all resources to be in a ready state before marking the deployment as successful. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: wait: true ``` **When to use:** - Ensure all Pods, PVCs, Services, and minimum number of Pods of Deployments/StatefulSets/ReplicaSets are ready - Guarantee that your application is fully operational before proceeding - Critical for production deployments where readiness is essential **Behavior:** When enabled, Helm will wait until all Kubernetes resources reach their ready state or until the timeout is reached. ### WaitForJobs **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--wait-for-jobs` Controls whether Helm waits for all Jobs to complete before marking the release as successful. This parameter only takes effect when `Wait` is also enabled. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: wait: true waitForJobs: true ``` **When to use:** - Applications that include initialization Jobs or batch processing - Database migrations or data seeding operations - Any scenario where Job completion is critical for application functionality **Note:** This parameter requires `Wait` to be set to `true` to function properly. ### TimeoutNanos **Type:** `uint64`\ **Default:** `300000000000` (5 minutes)\ **Helm Flag Equivalent:** `--timeout` Specifies the maximum time (in nanoseconds) to wait for any individual Kubernetes operation. Omnistrate stores this value in nanoseconds because the underlying runtime uses a duration type that is serialized in nanoseconds. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: timeoutNanos: 600000000000 # 10 minutes ``` **Common timeout values:** - 1 minute: `60000000000` - 5 minutes: `300000000000` (default) - 10 minutes: `600000000000` - 30 minutes: `1800000000000` **When to adjust:** - Large applications that take longer to start - Complex deployments with many dependencies - Environments with slower resource provisioning ### SkipCRDs **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--skip-crds` Controls whether Custom Resource Definitions (CRDs) are installed during the deployment. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: skipCRDs: true ``` **When to use:** - CRDs are already installed in the cluster - Managing CRDs separately from application deployments - Avoiding conflicts with existing CRD versions **Important:** Skipping CRDs when they're required by your application will cause deployment failures. ### UpgradeCRDs **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** Related to CRD upgrade behavior Determines whether CRDs should be upgraded during Helm operations. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: upgradeCRDs: true ``` **When to use:** - Ensuring CRDs are kept up-to-date with chart versions - Automated CRD lifecycle management - Charts that include CRD updates **Caution:** CRD upgrades can be disruptive and should be tested thoroughly in non-production environments. ### ResetValues **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--reset-values` When upgrading, resets the values to the ones built into the chart, ignoring any previously set values. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: resetValues: true ``` **When to use:** - Starting fresh with default chart values - Removing all custom configurations from previous deployments - Troubleshooting configuration-related issues **Warning:** This will remove all custom values from previous deployments. ### ReuseValues **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--reuse-values` When upgrading, reuses the last release's values and merges in any new overrides. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: reuseValues: true ``` **When to use:** - Preserving existing configuration during upgrades - Incremental configuration changes - Maintaining consistency across deployments **Note:** This is mutually exclusive with `ResetValues` and `ResetThenReuseValues`. ### ResetThenReuseValues **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--reset-then-reuse-values` When upgrading, first resets values to chart defaults, then applies the last release's values, and finally merges in any new overrides. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: resetThenReuseValues: true ``` **When to use:** - Ensuring clean configuration state while preserving custom values - Complex upgrade scenarios where value precedence matters - Resolving configuration conflicts **Behavior:** This provides a three-step process: reset → reuse → merge new values. ### Recreate **Type:** `bool`\ **Default:** `false`\ **Helm Flag Equivalent:** `--force` (similar behavior) Forces resource updates through a replacement strategy, recreating resources if they already exist. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: recreate: true ``` **When to use:** - Resolving resource conflicts or corruption - Forcing updates to immutable fields - Recovery from failed deployments `Recreate` does not uninstall the existing Helm release or wipe the namespace before reapplying the chart. It only forces replacement for resources that need it during the Helm operation. **Caution:** This can cause service interruption as resources are deleted and recreated. ### Disable Reconciliation **Type:** `bool`\ **Default:** `false` Disables steady-state reconciliation for an existing Helm release, preventing drift detection and correction after the release has been created. ``` services: - name: ExampleServiceName helmChartConfiguration: runtimeConfiguration: disableReconciliation: true ``` **When to use:** - Rely on the agent as a passive deployer only - Avoid automatic corrections to the deployed state - Scenarios where manual intervention is preferred - Testing or debugging deployments without steady-state reconciliation interference **Behavior notes:** - This setting applies after the release exists. Initial install and an in-flight Helm action can still retry until that workflow finishes. - It is not a blanket "never retry" switch for create or upgrade workflows. ## Complete Configuration Example ``` services: - name: My Application helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-repo chartRepoURL: https://charts.example.com helmRuntimeConfiguration: disableHooks: false wait: true waitForJobs: true timeoutNanos: 600000000000 # 10 minutes skipCRDs: false upgradeCRDs: true resetValues: false reuseValues: true resetThenReuseValues: false recreate: false disableReconciliation: false chartValues: app: name: "my-application" replicas: 3 ``` ## Troubleshooting ### Common Issues and Solutions **Deployment Timeouts:** - Increase `timeoutNanos` value - Check if `wait` and `waitForJobs` are necessary - Verify resource requests and limits **Hook Failures:** - Set `disableHooks: true` to bypass problematic hooks - Review hook configurations in your Helm chart **CRD Conflicts:** - Use `skipCRDs: true` if CRDs are managed separately - Enable `upgradeCRDs: true` for automatic CRD updates **Configuration Issues:** - Use `resetValues: true` to start with clean configuration - Enable `reuseValues: true` to preserve existing settings # Helm Chart and Terraform (or OpenTofu) ## How to configure Terraform (or OpenTofu) stacks Omnistrate supports integrating Terraform (or OpenTofu) IaC steps as part of a Helm chart deployment to build a unified deployment workflow. A good example of such a use-case is when you need auxiliary services like Databases or Caches through managed services like AWS RDS, Elasticache, MongoDB Atlas, etc. to be deployed as part of your Helm chart deployment. Additionally, you may want to reference the output of the Terraform IaC steps in the Helm chart configuration (e.g. the database connection string, cache endpoint, etc.). More details how to you can use Terraform (or OpenTofu) stack to define your service can be found on [Getting started / Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md). In this guide, we will show you how to deploy an Umbrella Helm chart that deploys a simple app that depends on a Postgres Helm chart with an S3 bucket and a DynamoDB table. You can define any number of such compositions and Omnistrate powers all of them through its idempotent, scalable workflow engine. The following diagram lays out the composition and the end goal for this SaaS deployment: At the end of this exercise, you will have an Nginx "app" deployment that is configured to talk to a Postgres database through a Kubernetes Service. The app also will have environment variables configured with the S3 bucket ARN and the DynamoDB Table ARN. To secure access to these services, we will leverage [IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html) for federated access. We will leverage the power and versatility of [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/) to inject the right context of the EKS cluster and OIDC configuration. ## Example: Hosted SaaS Deployment with Helm and Terraform (or OpenTofu) In this Hosted SaaS variant, we will deploy the entire composition for your tenants in your account. We will slice up your AWS account and EKS cluster into multiple tenant environments with isolation across the stack including dedicated IAM roles for each tenant deployment. ### Build a new SaaS Product plan / composition The first step is to model the above composition as a specification on Omnistrate. Following the earlier examples of onboarding a SaaS Product based on Helm charts, you can use the [Omnistrate CTL](https://docs.omnistrate.com/getting-started/getting-started-with-ctl/index.md) to build the SaaS Product plan. Make sure to replace the `` with your AWS account and [connect your AWS account](https://docs.omnistrate.com/getting-started/account-onboarding/index.md) to Omnistrate. ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json name: Web App # Plan Name deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: Web App dependsOn: - dataInfraTerraform compute: instanceTypes: - apiParam: instanceType cloudProvider: aws - apiParam: instanceType cloudProvider: gcp - apiParam: instanceType cloudProvider: azure network: ports: - 80 helmChartConfiguration: chartName: private-umbrella-chart chartVersion: 0.1.4 chartRepoName: private-umbrella-repo chartRepoURL: https://raw.githubusercontent.com/omnistrate-community/helm-private-example/main authProvider: username: password: chartValues: serviceAccount: name: "web-app-sa" s3BucketARN: "{{ $dataInfraTerraform.out.s3_bucket_arn }}" dynamoDBTableARN: "{{ $dataInfraTerraform.out.dynamodb_table_arn }}" affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname apiParameters: - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" - name: dataInfraTerraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: / gitConfiguration: reference: refs/heads/demo repositoryUrl: https://github.com/omnistrate-community/terraform-private-example.git accessToken: ``` For [**IntelliJ**](https://www.jetbrains.com/help/idea/yaml.html#use-schema-keyword) replace the top line with the following line to set up the yaml schema ``` # $schema: https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json ``` ### Apply changes to the service For this we will run the same command that was used to setup the service the first time. ``` omnistrate-ctl build -f spec.yaml --name 'UmbrellaHelm' --release-as-preferred --spec-type ServicePlanSpec # Example output shown below ✓ Successfully built service Check the Plan result at: https://omnistrate.cloud/product-tier?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm Access your SaaS Product at: https://saasportal.instance-w6vidhd14.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm ``` # Helm Chart Topology ## How to Configure Helm Charts In the [Getting Started](https://docs.omnistrate.com/getting-started/build-from-helm/index.md) guide, we designed a SaaS Product that deploys Redis Clusters using a Helm chart. However, in that example, we only deployed a single component as part of the overall SaaS deployment. Omnistrate supports deploying multiple Helm charts as part of a single deployment. Your SaaS Product may require multiple components to be deployed, and each component may have its own Helm chart. For eg: a Redis Cluster Helm chart and a Postgres Helm chart. In this guide, we will show you how to deploy multiple Helm charts as part of a single deployment. We will use the example of deploying a Redis Cluster and a Postgres database as part of a single deployment. ## Add the Helm Chart Configuration First, we need to update the Plan specification to include the Helm chart configurations for the Postgres Database. We will use the Bitnami chart to do so. We will also hide this component from your customers from a provisioning point of view and make it an internal component. We will also add some customization parameters that will be exposed to your customers. Info A complete description of the Plan specification can be found on [Getting started / Plan Spec](https://docs.omnistrate.com/spec-guides/plan-spec/index.md) ``` - name: Postgres Database internal: false compute: instanceTypes: - apiParam: postgresInstanceType cloudProvider: aws - apiParam: postgresInstanceType cloudProvider: gcp - apiParam: postgresInstanceType cloudProvider: azure network: ports: - 5432 helmChartConfiguration: chartName: postgresql chartVersion: 15.5.36 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartValues: auth: database: $var.postgresDatabase username: $var.postgresUsername password: $var.postgresPassword primary: persistence: enabled: false resources: requests: cpu: 100m memory: 128Mi limits: cpu: 1000m memory: 1024Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname readReplicas: replicaCount: 0 apiParameters: - key: postgresUsername description: Postgres Username name: Postgres Username type: String modifiable: true required: true export: true - key: postgresPassword description: Postgres Password name: Postgres Password type: String modifiable: true required: true export: true - key: postgresDatabase description: Postgres Database name: Postgres Database type: String modifiable: true required: true export: true - key: postgresInstanceType description: Postgres Instance Type name: Postgres Instance Type type: String modifiable: true required: true export: true ``` We will then set the Postgres component as a dependency on the Redis component. This will ensure that the Postgres component is deployed before the Redis component. ``` services: - name: Redis Cluster dependsOn: - Postgres Database apiParameters: ... - key: postgresUsername description: Postgres Username name: Postgres Username type: String modifiable: false required: false export: true defaultValue: "postgres" parameterDependencyMap: Postgres Database: postgresUsername - key: postgresPassword description: Postgres Password name: Postgres Password type: Password modifiable: false required: true export: true parameterDependencyMap: Postgres Database: postgresPassword - key: postgresDatabase description: Postgres Database name: Postgres Database type: String modifiable: false required: false export: true defaultValue: "postgres" parameterDependencyMap: Postgres Database: postgresDatabase - key: postgresInstanceType description: Postgres Instance Type name: Postgres Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" parameterDependencyMap: Postgres Database: postgresInstanceType ``` You will also notice we copied over the new API parameters from the Postgres component to the Redis component. The platform requires any dependent component to provide the necessary parameters to the child component as a pass-through. This is done by setting the `parameterDependencyMap` field in the API parameters and mapping them to the dependent component and the relevant API parameter defined on the dependent component. ## Helm Chart specification example ``` name: Redis Server # Plan Name deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: Redis Cluster dependsOn: - Postgres Database compute: instanceTypes: - apiParam: instanceType cloudProvider: aws - apiParam: instanceType cloudProvider: gcp - apiParam: instanceType cloudProvider: azure network: ports: - 6379 helmChartConfiguration: chartName: redis chartVersion: 24.1.0 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartValues: master: podLabels: omnistrate.com/schedule-mode: exclusive persistence: enabled: false service: type: LoadBalancer annotations: external-dns.alpha.kubernetes.io/hostname: $sys.network.externalClusterEndpoint resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname replica: podLabels: omnistrate.com/schedule-mode: exclusive persistence: enabled: false replicaCount: $var.replicas resources: requests: cpu: 100m memory: 128Mi limits: cpu: 150m memory: 256Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname apiParameters: - key: replicas description: Number of Replicas name: Replica Count type: Float64 modifiable: true required: false export: true defaultValue: "1" - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" - key: postgresUsername description: Postgres Username name: Postgres Username type: String modifiable: false required: false export: true defaultValue: "postgres" parameterDependencyMap: Postgres Database: postgresUsername - key: postgresPassword description: Postgres Password name: Postgres Password type: Password modifiable: false required: true export: true parameterDependencyMap: Postgres Database: postgresPassword - key: postgresDatabase description: Postgres Database name: Postgres Database type: String modifiable: false required: false export: true defaultValue: "postgres" parameterDependencyMap: Postgres Database: postgresDatabase - key: postgresInstanceType description: Postgres Instance Type name: Postgres Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" parameterDependencyMap: Postgres Database: postgresInstanceType - name: Postgres Database internal: false compute: instanceTypes: - apiParam: postgresInstanceType cloudProvider: aws - apiParam: postgresInstanceType cloudProvider: gcp - apiParam: postgresInstanceType cloudProvider: azure network: ports: - 5432 helmChartConfiguration: chartName: postgresql chartVersion: 15.5.36 chartRepoName: bitnami chartRepoURL: https://charts.bitnami.com/bitnami chartValues: auth: database: $var.postgresDatabase username: $var.postgresUsername password: $var.postgresPassword primary: persistence: enabled: false resources: requests: cpu: 100m memory: 128Mi limits: cpu: 1000m memory: 1024Mi affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/managed-by operator: In values: - omnistrate - key: topology.kubernetes.io/region operator: In values: - $sys.deploymentCell.region - key: node.kubernetes.io/instance-type operator: In values: - $sys.compute.node.instanceType - key: omnistrate.com/resource operator: In values: - $sys.deployment.resourceID podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: omnistrate.com/schedule-mode operator: In values: - exclusive namespaceSelector: {} topologyKey: kubernetes.io/hostname readReplicas: replicaCount: 0 apiParameters: - key: postgresUsername description: Postgres Username name: Postgres Username type: String modifiable: true required: true export: true - key: postgresPassword description: Postgres Password name: Postgres Password type: String modifiable: true required: true export: true - key: postgresDatabase description: Postgres Database name: Postgres Database type: String modifiable: true required: true export: true - key: postgresInstanceType description: Postgres Instance Type name: Postgres Instance Type type: String modifiable: true required: true export: true ``` ### Apply changes to the service For this we will run the same command that was used to setup the service the first time. ``` omnistrate-ctl build -f spec.yaml --name 'RedisHelm' --release-as-preferred --spec-type ServicePlanSpec # Example output shown below ✓ Successfully built service Check the Plan result at: https://omnistrate.cloud/product-tier?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm Access your SaaS Product at: https://saasportal.instance-w6vidhd14.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud/service-plans?serviceId=s-dEhutaDa2X&environmentId=se-92smpU2YAm ``` ### Deploying a Redis Cluster through your dedicated Customer Portal Now, your customers can deploy Redis Clusters alongwith the Postgres Helm chart with the desired instance type and number of replicas through the dedicated customer portal. The portal will use the REST API parameters to customize the deployment based on the customer's requirements. The deployed cluster details The Kubernetes Dashboard view # Helm Charts Troubleshooting For the shared debugging workflow, start with [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md). It explains Debug Events, instance debug, and the relationship between workflow state and live resource state. ## Helm Troubleshooting Checklist If a Helm-based deployment fails or stays in a provisioning state, start from [Debug Events](https://docs.omnistrate.com/operate-guides/troubleshooting/#debug-events) to identify the failed stage and resource. Then inspect the Helm resource with: ``` omnistrate-ctl instance debug ``` Use the instance debug output to investigate: 1. Whether the rendered chart values match the API parameters, system parameters, and chart defaults you expected. 1. Helm client logs for template rendering errors, hook failures, timeout messages, or Kubernetes API validation errors. 1. Kubernetes resources created by the chart, especially `Jobs`, hooks, `Services`, load balancers, readiness probes, and pending pods. 1. Leftover CRDs, finalizers, or namespaced resources if create, upgrade, or delete workflows repeatedly fail. 1. Runtime settings such as `wait`, `waitForJobs`, `disableHooks`, `skipCRDs`, `upgradeCRDs`, and `timeoutNanos` if the release lifecycle does not match the chart behavior. The shared [Helm Troubleshooting Checklist](https://docs.omnistrate.com/operate-guides/troubleshooting/#helm-troubleshooting-checklist) contains the same checks alongside the other deployment strategies. If you change chart values, API parameter mappings, chart source, or other rendered artifacts, publish a new Plan version and trigger a fresh workflow instead of repeatedly restarting the same failed workflow. For runtime flag details, see [Helm Runtime Configuration](https://docs.omnistrate.com/build-guides/helm-charts-runtime-configuration/index.md). # Integrations Overview Integrations are 1st/3rd party SaaS applications that you want to integrate with for your SaaS. We have classified SaaS integrations into the following categories: - **Observability**: Examples - Datadog, NewRelic, Loggly, Signoz, Axiom - **Billing**: Examples - Stripe - **Continuous Integration**: Examples - GitHub Actions, CircleCI - **Marketplace**: Examples - Clazar, Suger, Feenix, Labra, Tackle - **Cloud Insurance**: Example - Archera - **Operational Status**: Examples - Atlassian, BetterStack If you don't see a category for your SaaS integration and would like us to support, please reach out to us [support@omnistrate.com](mailto:support@omnistrate.com) # Observability integrations Omnistrate supports 2 different 'scopes' for observability features: - Internal - provides visibility for SaaS Provider across fleet - Customer - provides visibility for your customer into their instance(s) Different observability providers are available based on intended scope. # Omnistrate native Omnistrate internal observability creates dashboards accessible by SaaS Provider on fleet page. You can enable logging dashboard integration from your compose file using: ``` x-internal-integrations: logs: ``` Enabling this integration will enable real-time logging for your support team to manage the customer fleet. You can enable metrics dashboard integration from your compose file using: ``` x-customer-integrations: metrics: ``` Enabling this integration will enable real-time infra metrics dashboard. Note We launch a service deployment in same kubernetes cluster where instances are running, this deployment is responsible to collect, store, aggregate, and stream metrics and logs across all nodes running in the same kubernetes cluster. ## Custom metrics and dashboards (advanced) The advanced metrics model has three layers: 1. **Collection layer** (`scrapeTargets`): define where Prometheus should scrape from. 1. **Series modeling layer** (`additionalMetrics..metrics`): transform raw metric names into stable, human-readable series with optional filtering and aggregation. 1. **Visualization layer** (`additionalMetrics..dashboards`): define custom dashboard layout (sections, panels, targets) using modeled series names. For Helm/Operator/Kustomize/Terraform plans (service `spec.yaml`), use `features.CUSTOMER.metrics` or `features.INTERNAL.metrics`. For Compose specs, use the equivalent `x-customer-integrations.metrics` or `x-internal-integrations.metrics` blocks. ``` features: CUSTOMER: metrics: scrapeTargets: - name: postgres podSelector: app.kubernetes.io/name: postgresql app.kubernetes.io/instance: postgres endpoint: portNumber: 9187 path: /metrics scheme: http additionalMetrics: postgres: metrics: pg_stat_activity_count: "Active Sessions": aggregationFunction: sum labelFilters: datname: postgres state: active pg_locks_count: "Locks by Mode": aggregationFunction: sum groupByLabels: - mode labelFilters: datname: postgres dashboards: exporter: title: "Postgres Exporter" sections: - title: "Sessions and Locks" panels: - title: "Active Sessions" type: timeseries unit: short targets: - ref: "Active Sessions" - title: "Locks by Mode" type: timeseries unit: short targets: - ref: "Locks by Mode" ``` - **`scrapeTargets`** decides what gets ingested at all. If a target is missing here, `additionalMetrics` cannot resolve it. - **`metrics` map keys** are raw Prometheus metric names; nested keys are your exported series names used by dashboards. - **`labelFilters`** narrow to business-relevant series (for example, one database or state). - **`aggregationFunction`** controls how multiple matched series are reduced (`sum`, `avg`, `max`, `min`). - **`groupByLabels`** preserves selected dimensions (for example, `mode`) instead of collapsing everything into one line. - **`dashboards`** should reference modeled series via `targets[].ref`, which keeps panel definitions stable even if raw Prometheus labels evolve. Panel capabilities in dashboards: - `type`: `timeseries`, `stat`, `gauge`, `table` - `targets`: use `ref` (modeled series) or `expr` (raw PromQL), optional `legend` and `axis` - optional presentation fields: `description`, `unit`, `min`, `max`, `thresholds`, `tags` Dashboard property reference: | Path | Type | Required | What it controls | Guidance | | ------------------------------------ | ------------- | ------------------------ | ---------------------------------------------------------------------------- | ----------------------------------------------------------------- | | `dashboards..title` | string | No | Dashboard title shown in Grafana | Use a product-friendly title, for example `Postgres Exporter` | | `dashboards..sections` | array | Yes | Ordered sections in the dashboard | Keep sections task-oriented (Health, Sessions, Replication, etc.) | | `sections[].title` | string | Yes | Section header | Use concise titles so users can scan quickly | | `sections[].description` | string | No | Section-level help text | Use for context or runbook hints | | `sections[].panels` | array | Yes | Panels rendered in that section | Group related metrics only | | `panels[].title` | string | Yes | Panel title | Prefer metric intent over metric name | | `panels[].type` | string enum | Yes | Visualization type | Allowed: `timeseries`, `stat`, `gauge`, `table` | | `panels[].description` | string | No | Panel-level explanatory text | Add when panel logic is non-obvious | | `panels[].unit` | string | No | Value unit format | Examples: `short`, `bytes`, `s`, `%` | | `panels[].min` / `panels[].max` | number | No | Fixed panel value range | Useful for `gauge` and bounded KPIs | | `panels[].targets` | array | Yes | Data queries/series in panel | Provide at least one target per panel | | `targets[].ref` | string | Recommended | Reference to modeled series name from `additionalMetrics..metrics` | Preferred for stability across metric label changes | | `targets[].expr` | string | Optional alternative | Raw PromQL query | Use only when `ref` cannot express needed query | | `targets[].legend` | string | No | Series legend label | Improves readability in multi-series panels | | `targets[].axis` | string enum | No | Axis selection for mixed-unit panels | Use `left`/`right` when combining heterogeneous units | | `panels[].thresholds` | array | No | Threshold markers and colors | Good for `stat`/`gauge` health indicators | | `thresholds[].color` | string | Yes (if thresholds used) | Threshold color | Typical values: `green`, `yellow`, `red` | | `thresholds[].value` | number | Yes (if thresholds used) | Threshold boundary | Order thresholds from lower to higher severity | | `panels[].tags` | array[string] | No | Panel tags | Use for filtering/grouping conventions | Recommended target strategy: - Use `ref` for most panels so dashboards remain resilient to raw metric/schema changes. - Use `expr` only for derived or cross-metric PromQL that cannot be represented by one modeled series. Panel design patterns (copy/paste templates): 1. Health status (`stat` + thresholds) ``` - title: "Exporter Up" type: stat unit: short thresholds: - color: red value: 0 - color: green value: 1 targets: - ref: "Postgres Up" ``` Use this for binary states and quick health checks. 1. Saturation trend (`timeseries` + optional dual axis) ``` - title: "Connections and Max Capacity" type: timeseries unit: short targets: - ref: "DB Connections" legend: "connections" axis: left - ref: "DB Max Connections" legend: "max" axis: right ``` Use this when you want usage vs capacity in one panel. If you need utilization %, add a dedicated `expr` target with raw PromQL. 1. Distribution breakdown (`table` or `timeseries` with `groupByLabels`) ``` metrics: pg_locks_count: "Locks by Mode": aggregationFunction: sum groupByLabels: - mode labelFilters: datname: postgres dashboards: exporter: sections: - title: "Locks" panels: - title: "Locks by Mode" type: table targets: - ref: "Locks by Mode" ``` Use this for category distributions (for example lock mode, status, operation type). ## Migration guide: legacy custom metrics to advanced metrics + dashboards If you previously configured only exported metrics and now want custom dashboards, migrate in this order: 1. Move endpoint discovery to `scrapeTargets`. 1. Keep/refine metric modeling in `additionalMetrics..metrics`. 1. Add `additionalMetrics..dashboards` and wire panels via `targets.ref`. Legacy to new field mapping: | Legacy pattern | New pattern | Why | | ------------------------------------------------- | ---------------------------------------------------------- | --------------------------------------------------------- | | `additionalMetrics..prometheusEndpoint` | `metrics.scrapeTargets[].endpoint` + selector | Separates scrape discovery from metric modeling | | `additionalMetrics..metrics` | `additionalMetrics..metrics` (unchanged concept) | Keeps named, stable modeled series | | no dashboard block | `additionalMetrics..dashboards` | Enables first-class custom layouts and panel semantics | | panel logic implicit in default dashboards | explicit `sections[].panels[].targets` | Makes visualization deterministic and versionable in spec | Before (legacy minimal): ``` x-customer-integrations: metrics: additionalMetrics: postgres: prometheusEndpoint: "http://localhost:9187/metrics" metrics: pg_stat_activity_count: active_sessions: aggregationFunction: sum labelFilters: state: active ``` After (advanced): ``` x-customer-integrations: metrics: scrapeTargets: - name: postgres podSelector: app.kubernetes.io/name: postgresql endpoint: portNumber: 9187 path: /metrics scheme: http additionalMetrics: postgres: metrics: pg_stat_activity_count: "Active Sessions": aggregationFunction: sum labelFilters: state: active dashboards: exporter: title: "Postgres Exporter" sections: - title: "Activity" panels: - title: "Active Sessions" type: timeseries targets: - ref: "Active Sessions" ``` Validation checklist during migration: - Ensure each scrape source has a matching `scrapeTargets` selector + endpoint. - Keep modeled series names stable; dashboards depend on `targets.ref` exact matches. - Apply `labelFilters` before adding aggregation to avoid mixing dimensions. - Start with one dashboard section and 2-3 panels, then expand iteratively. Warning `additionalMetrics` and custom dashboards are supported for Omnistrate native observability. For 3rd party OTEL providers, use `serviceComponentsConfiguration` to define scrape endpoints and build visualization in your OTEL backend. # OpenTelemetry providers Supported open-telemetry providers: NewRelic, Signoz, Datadog Omnistrate allows you to ship metrics to your open telemetry provider account. This will enable you to: - Provide visibility to your customers. - Stage/debug any form of production issues that your customers may be facing while provisioning/operating your SaaS. ## Metrics You can enable metrics integration by following the steps below: 1. Generate an API key from your open telemetry providers account and store them in the cloud native secret manager (for AWS that would be "Secrets Manager"). 1. You can then enable metrics integration by adding the following to your compose specification: ``` x-internal-integrations: metrics: provider: endpoint: secretLocators: gcp: projects/xxxxxxxxxxx/secrets/mySecret123456-abc123/versions/latest aws: arn:aws:secretsmanager:us-west-2:xxxxxxxxxxx:secret:mySecret123456-abc123 azure: KeyVaultName/SecretName ``` The specification includes the following parameters: 1. Provider name - specifies which open-telemetry provider to use. Supported values: - `signoz` - `newRelic` - `datadog` 1. Provider endpoint - specify endpoint for open telemetry provider of choice. Examples: - `ingest.us.signoz.cloud:443` - for Signoz. For list of available cloud endpoints, see [Signoz docs](https://signoz.io/docs/ingestion/signoz-cloud/overview/). You can also export OTEL telemetry to self-hosted Signoz instance. - `https://otlp.nr-data.net` - for NewRelic. For list of available endpoints, see [NewRelic docs](https://docs.newrelic.com/docs/opentelemetry/best-practices/opentelemetry-otlp/). - `us5.datadoghq.com` - for Datadog. This endpoint needs to be one of the datadog ingest endpoints. For more information, see one of the options named "Site Parameter" in [Datadog docs](https://docs.datadoghq.com/getting_started/site/#access-the-datadog-site). 1. Providing secret store locator for API key - specify secret identifier from within secret store of used cloud provider. - For AWS, we expect ARN of secret from within "AWS Secrets Manager" - For GCP, we expect secret ID (including version) of secret stored in "Secret Manager". - For Azure, we expect identifier in form of `KeyVaultName/SecretName` where `KeyVaultName` is the name of the "Azure Key Vault" and `SecretName` is the name of the secret within this vault. In order to be able to use this secret, it needs to share Azure Subscription with the target account config for deployment. Note Most otel providers require single API key / token as authentication mechanism. This "key" can be directly stored as secret value and would be used as needed based on the otel provider. In some scenarios, such as self-hosting or hosting Signoz using Omnistrate, authentication is performed using combination of username and password. In that case, secret value needs to be in form of following json object that provides both: `{"username":"replace_with_username","password":"replace_with_password"}` We will automatically add infrastructure metrics. If you want to add your own custom metrics in addition, you can define your own prometheus endpoint (per Resource) for us to scrape and add them. You can enable custom metrics by adding the following to your Resource: ``` metrics: serviceComponentsConfiguration: postgres: # Resource prometheusEndpoint: "http://localhost:9187/metrics" # Add the local host endpoint to supply promotheus metrics. ``` Omnistrate will automatically scrape the metrics of the services and ship them to your open telemetry provider account. ## Logging Logging can be enabled using same configuration parameters, but specifying 'logs' as feature: ``` x-internal-integrations: logs: provider: endpoint: secretLocators: gcp: projects/xxxxxxxxxxx/secrets/mySecret123456-abc123/versions/latest aws: arn:aws:secretsmanager:us-west-2:xxxxxxxxxxx:secret:mySecret123456-abc123 azure: KeyVaultName/SecretName ``` Logging integration will upload service logs to open telemetry provider of your choice. Both metrics/logs can be activated using API method with following format: ``` { "feature": "METRICS", "scope": "INTERNAL", "configuration": { "provider": "newRelic", "endpoint": "https://otlp.nr-data.net", "secretLocators": { "aws": "arn:aws:secretsmanager:us-west-2:xxxxxxxxxxx:secret:mySecret123456-abc123", "gcp": "projects/my-project/secrets/my-secret/versions/latest", "azure": "KeyVaultName/SecretName" }, "serviceComponentsConfiguration": { "postgres": { "prometheusEndpoint": ":9090" } } } } ``` Note - Logs use "LOGS" as feature name and will not expect "serviceComponentsConfiguration" parameter. - For metrics, "serviceComponentsConfiguration" parameter with additional prometheus endpoint configuration is optional. Info Please note that this integration is only available in the 'Hosted SaaS' Deployment Model. ## Cloud native observability Your customers' observability is also supported when using the BYOC model. When your customers deploy in their own cloud account, they get native observability through their cloud provider's built-in platform (CloudWatch, Azure Monitor, or Cloud Operations Suite). Activate it in similar fashion as internal native observability using your Compose specification: ``` x-customer-integrations: logs: provider: native metrics: provider: native ``` # Custom OpenTelemetry In case an open-telemetry provider of your choice is not supported out of the box, you can configure your own sidecar with an open-telemetry collector of your choice. Please see the [Using own OTEL exporter](https://github.com/omnistrate-community/examples/blob/main/docs/custom-otel-sidecar/README.md) for an example how to easily do it with [opentelemetry-collector-contrib](https://github.com/open-telemetry/opentelemetry-collector-contrib) container. ## Cloud native Each cloud provider has its own built-in observability platform. Omnistrate automatically configures the required permissions and integrations for the cloud provider where your deployment runs. | Cloud Provider | Logs | Metrics | Native Platform | | -------------- | --------- | --------- | -------------------- | | AWS | Supported | Supported | Amazon CloudWatch | | Azure | Supported | Supported | Application Insights | | GCP | Supported | Supported | Operations Suite | Note For Azure deployments, Omnistrate configures the necessary service principal and role assignments to push logs and metrics to Azure Monitor. No additional setup is required beyond enabling the integration in your compose specification. To enable native integration from the compose specification, add the following configuration for logs: ``` x-internal-integrations: logs: provider: native ``` For metrics: ``` x-internal-integrations: metrics: provider: native ``` ## Omnistrate Metering Enabling this integration will enable metering from your infrastructure to capture the usage per customer, aggregate and store them at a defined location in your account. To enable, please see [here](https://docs.omnistrate.com/fin-ops-guides/metering/index.md) ## Omnistrate Billing If you want to charge your customers directly, you can enable billing integration. To enable, please see [here](https://docs.omnistrate.com/fin-ops-guides/billing/index.md) ## Alarming Omnistrate platform allows you to enable real-time alarming with PagerDuty and other services. Omnistrate will notify you when there is an issue with your SaaS infrastructure and enabling your operators to take action. For example, we can send an alert when your deployments fail or are not available. To configure Alerts please see [here](https://docs.omnistrate.com/operate-guides/alarms/index.md). ## Marketplace If you want to list your offerings to cloud marketplace, you can enable marketplace integration. Pelase review the Marketplace integration guideline [here](https://docs.omnistrate.com/fin-ops-guides/marketplace/index.md) Info Please note that this integration is only available in the enterprise plan. ## Cloud Insurance If you want to purchase cloud with unprecedented flexibility, you can enable cloud insurance integration. Info Please note that this integration is only available in the enterprise plan. ## Operational Status If you want to want to keep your customers informed about service status and streamline incident reporting, you can enable this integration by reaching out to us at [support@omnistrate.com](mailto:support@omnistrate.com). Info Please note that this integration is only available in the enterprise plan for now. ## Future Integrations We are constantly adding native support for new integrations. If you would like us to prioritize any particular integration, please reach out us at [support@omnistrate.com](mailto:support@omnistrate.com) Note Please note that you can also directly integrate your favorite tooling yourself as you have full control on the underlying infrastructure If you looking to partner and integrate your solution with us, please reach out to us at [partnerships@omnistrate.com](mailto:partnerships@omnistrate.com) # Kubernetes Operators Deployment Strategy The Kubernetes Operators strategy lets you keep an operator as the application lifecycle layer while Omnistrate supplies the control plane around it. The operator reconciles custom resources; Omnistrate manages the product, tenant, infrastructure, and operational surfaces needed to deliver that application as a service. ## When to Use Kubernetes Operators The Operator strategy is a good fit when: - Your product lifecycle is already modeled by Kubernetes CRDs and an operator. - The application needs domain-specific reconciliation, upgrades, backup and restore, failover, or status handling. - You want customer-facing APIs and portals without replacing the operator’s application logic. - You need to expose standard lifecycle operations together with provider-defined operations. ## Responsibility Boundary The operator owns application-specific state: CRDs, reconciliation, workload topology, application health, and domain-specific recovery or upgrade behavior. Omnistrate owns the surrounding service lifecycle: tenants, subscriptions, APIs, infrastructure, deployment orchestration, versioned releases, observability, metering, billing, and multi-cloud or BYOC delivery. This separation keeps the operator focused on running the application while the control plane provides a consistent experience across customers and deployment environments. ## Strategy Guide - Start with [Build from Kubernetes Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md) for the short onboarding path. - Use [Build with Kubernetes Operators](https://docs.omnistrate.com/build-guides/operators/index.md) for service-plan modeling, system workflows, custom workflows, and backup/restore behavior. - See [Operator Troubleshooting](https://docs.omnistrate.com/build-guides/operators-troubleshooting/index.md) when an operator-backed deployment or workflow fails. - Review the [generic troubleshooting workflow](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md) for shared debugging tools and instance-debug behavior. For the enterprise control-plane perspective, see [Beyond the Operator: The Enterprise Control Plane Layer](https://omnistrate.com/blog/beyond-operators-the-enterprise-control-plane-layer). # Kubernetes Operator Troubleshooting For the shared debugging workflow, start with [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md). It explains Debug Events, instance debug, and the relationship between workflow state and live resource state. ## Operator Troubleshooting Checklist Use `omnistrate-ctl instance debug ` to open the Operator resource and work through the following: 1. Review application logs from the operator-managed workload. 1. Confirm deployment API parameters resolved to the expected custom-resource inputs. 1. Inspect Operator CRD outputs and exported deployment output parameters. 1. Check workflow events to separate Omnistrate orchestration errors from operator reconciliation errors. 1. Use the Metrics tab or `omnistrate-ctl instance dashboard ` to verify dashboard access when metrics are enabled. 1. Use [Deployment Cell Access](https://docs.omnistrate.com/operate-guides/deployment-cell-access/index.md) to inspect the custom resource, operator logs, events, and status conditions directly. The shared [Operator Troubleshooting Checklist](https://docs.omnistrate.com/operate-guides/troubleshooting/#operator-troubleshooting-checklist) contains the same operational checks alongside the Compose, Helm, and Terraform checklists. ## Workflow and Resource Failures Check the operator’s status fields and the `successCondition` and `failureCondition` expressions used by the workflow. If readiness stalls or an output is not resolved, confirm that the operator writes the referenced status path and that the workflow targets the intended namespace and resource name. # Build with Kubernetes Operators Use this guide when your product is already managed by a Kubernetes Operator and you want Omnistrate to turn it into a customer-facing SaaS Product. Operators are a good fit when your application lifecycle is already expressed as Kubernetes custom resources: database clusters, message queues, streaming platforms, storage systems, AI platforms, or any product where an operator reconciles the desired state. Omnistrate keeps that operator model intact and adds the control plane around it: cloud account onboarding, tenant management, APIs, Customer Portal, subscriptions, deployment cells, networking, backups, restores, upgrades, observability, and fleet operations. For a complete working example, use the [operator spec template](https://github.com/omnistrate-community/operator-spec-template). It builds a CloudNativePG-based PostgreSQL service plan with create, modify, start, stop, scale, backup, restore, and delete-backup system workflows. ## What You Define An operator-backed service plan usually has four layers: | Layer | Where it is defined | Purpose | | --------------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | Customer inputs | `apiParameters` | Values your customers provide, such as instance type, replica count, storage size, database name, or backup bucket. | | Operator installation | `operatorCRDConfiguration.helmChartDependencies` or deployment-cell amenities | Installs the operator and required CRDs before tenant resources are reconciled. | | Runtime lifecycle | `systemWorkflows` | Uses Argo Workflow-style DAGs to create, update, delete, start, stop, scale, back up, restore, and delete backups. | | Provider operations | Optional `customWorkflows` | Exposes provider-defined actions that are not platform lifecycle APIs, such as compact, repair, diagnostics, or product-specific administrative tasks. | Note `operatorCRDConfiguration.template`, `operatorCRDConfiguration.supplementalFiles`, and `operatorCRDConfiguration.readinessConditions` are deprecated for operator lifecycle management. Prefer `systemWorkflows` because they let you model multi-step lifecycle operations, capture outputs, run backups and restores, and use Kubernetes resource success and failure conditions with the same workflow engine. Keep `operatorCRDConfiguration.helmChartDependencies` when the service plan should install the operator and CRDs. ## Minimal Service Plan Shape The operator spec template starts with a normal service plan: ``` name: Postgres Operator deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: "arn:aws:iam:::role/omnistrate-bootstrap-role" tenancyType: CUSTOM_TENANCY services: - name: CNPG compute: instanceTypes: - apiParam: instanceType cloudProvider: aws apiParameters: - key: instanceType name: Instance Type type: String modifiable: true defaultValue: "t3.medium" - key: numberOfInstances name: Total Number of Instances type: Float64 modifiable: true defaultValue: "1" - key: storageSize name: Storage Size type: String modifiable: true defaultValue: "20Gi" ``` Add endpoint configuration for customer-facing connection details: ``` endpointConfiguration: writer: host: "$sys.network.externalClusterEndpoint" ports: - 5432 primary: true networkingType: PUBLIC reader: host: "reader-{{ $sys.network.externalClusterEndpoint }}" ports: - 5432 primary: false networkingType: PUBLIC ``` Install the operator as a Helm dependency when the operator version should be tied to the service plan version: ``` operatorCRDConfiguration: helmChartDependencies: - chartName: cloudnative-pg chartVersion: 0.28.2 chartRepoName: cnpg chartRepoURL: https://cloudnative-pg.github.io/charts - chartName: plugin-barman-cloud chartVersion: 0.6.0 chartRepoName: cnpg chartRepoURL: https://cloudnative-pg.github.io/charts ``` Note The current operator template intentionally omits the deprecated `operatorCRDConfiguration.template`, `supplementalFiles`, and `readinessConditions` fields. Lifecycle resources, readiness, and outputs are modeled in `systemWorkflows` instead. If the operator is shared by many tenant instances in the same deployment cell, install it as a [deployment-cell amenity](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) instead. That lets you upgrade the operator once per cluster instead of coupling it to every tenant instance upgrade. ## Backup Capability Backup policy belongs to the resource capability, not the workflow definition: ``` capabilities: backupConfiguration: backupRetentionInDays: 1 backupPeriodInHours: 1 snapshotBeforeDeletion: true ``` Omnistrate uses this configuration to schedule automated backups, set expiration for automated backups, run backup-before-delete when enabled, and trigger delete-backup cleanup for expired snapshots. Manual snapshots are user-controlled and are not expired by this schedule. ## System Workflows `systemWorkflows` are lifecycle hooks invoked by Omnistrate's existing platform APIs. They use an Argo Workflow-style structure: `entrypoint`, `arguments.parameters`, `templates`, DAG `tasks`, and Kubernetes `resource` templates. The operator template uses the following system workflows: | Workflow | Trigger | Example behavior | | -------------- | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | `create` | Instance provisioning | Create secrets, create backup object-store configuration, apply the CNPG `Cluster`, and wait for ready instances. | | `modify` | Instance update or upgrade | Reapply the `Cluster` with updated inputs such as storage or replica count. | | `delete` | Instance delete | Delete the `Cluster`, secrets, and backup object-store resources. | | `start` | Start API | Patch the CNPG hibernation annotation to `off`. | | `stop` | Stop API | Patch the CNPG hibernation annotation to `on`. | | `backup` | Manual backup, periodic backup, backup-before-delete | Rehydrate the cluster if needed, create a CNPG `Backup` CR, and persist backup metadata. | | `restore` | Restore API | Create a new target cluster from selected snapshot metadata. | | `deleteBackup` | Manual snapshot delete and retention cleanup | Delete the operator backup CR or external backup marker. | At minimum, operator-backed service plans should define `create`, `modify`, and `delete` so Omnistrate can provision, update, and remove the managed resource through standard lifecycle APIs. Add the other system workflows only for the operations your service plan supports. ### Create workflow example This simplified example shows the shape of a `create` workflow. The full template includes secrets, backup object stores, load balancer annotations, affinity, and output parameters. ``` systemWorkflows: create: outputParameters: postgresContainerImage: "$tasks.applycluster.resource.status.image" status: "$tasks.applycluster.resource.status.phase" topology: "$tasks.applycluster.resource.status.topology" workflow: entrypoint: create arguments: parameters: - name: namespace value: "{{ $sys.namespace }}" - name: instanceId value: "{{ $sys.instanceId }}" - name: numberOfInstances value: "{{ $var.numberOfInstances }}" - name: storageSize value: "{{ $var.storageSize }}" templates: - name: create inputs: parameters: - name: namespace - name: instanceId - name: numberOfInstances - name: storageSize dag: tasks: - name: applycluster template: apply-cluster arguments: parameters: - name: namespace value: "{{inputs.parameters.namespace}}" - name: instanceId value: "{{inputs.parameters.instanceId}}" - name: numberOfInstances value: "{{inputs.parameters.numberOfInstances}}" - name: storageSize value: "{{inputs.parameters.storageSize}}" - name: apply-cluster inputs: parameters: - name: namespace - name: instanceId - name: numberOfInstances - name: storageSize resource: action: apply successCondition: status.instances == {{inputs.parameters.numberOfInstances}} && status.readyInstances == {{inputs.parameters.numberOfInstances}} failureCondition: status.phase == failed manifest: | apiVersion: postgresql.cnpg.io/v1 kind: Cluster metadata: name: "{{inputs.parameters.instanceId}}" namespace: "{{inputs.parameters.namespace}}" spec: instances: {{inputs.parameters.numberOfInstances}} storage: size: "{{inputs.parameters.storageSize}}" storageClass: gp3 ``` The `resource.action` can be `apply`, `patch`, or `delete`. `successCondition` and `failureCondition` are evaluated against the live Kubernetes resource status. Output parameters can reference completed task resources through `$tasks..resource.*`. ## Backup and Restore Workflow Context Backup workflows receive Omnistrate snapshot context: ``` systemWorkflows: backup: outputParameters: backupId: "$tasks.createBackupCR.resource.status.backupId" backupName: "$tasks.createBackupCR.resource.status.backupName" workflow: arguments: parameters: - name: snapshotId value: "{{ $sys.snapshot.id }}" - name: snapshotTime value: "{{ $sys.snapshot.time }}" - name: namespace value: "{{ $sys.namespace }}" - name: instanceId value: "{{ $sys.instanceId }}" ``` Values rendered from `backup.outputParameters` are stored as snapshot metadata. Restore workflows can use that metadata: ``` systemWorkflows: restore: workflow: arguments: parameters: - name: restoreSnapshotId value: "{{ $sys.restore.snapshotId }}" - name: restoreSnapshotTime value: "{{ $sys.restore.snapshotTime }}" - name: backupId value: "{{ $sys.restore.metadata.backupId }}" - name: backupName value: "{{ $sys.restore.metadata.backupName }}" - name: sourceInstanceId value: "{{ $sys.sourceInstanceId }}" - name: newClusterName value: "{{ $sys.targetInstanceId }}" ``` This keeps provider backup identifiers out of the customer restore form. Customers select an Omnistrate snapshot; the restore workflow receives the metadata captured during the original backup. ## Custom Workflows Use `customWorkflows` for provider-defined operations that are not platform lifecycle APIs. The operator template does not include a custom workflow by default; it focuses on the standard lifecycle operations implemented as `systemWorkflows`. Add custom workflows only when your product needs an additional operation beyond create, modify, delete, start, stop, scale, backup, restore, or delete-backup. Custom workflows appear in `supportedOperations` for the instance and can be invoked from the UI, API, or `omnistrate-ctl` operation commands. ## Argo Workflow Syntax The workflow body follows the Argo Workflow model for DAGs and Kubernetes resource templates. The most important concepts are: - `entrypoint` selects the template to run first. - `arguments.parameters` binds Omnistrate variables to workflow inputs. - `templates[].dag.tasks` declares the task graph and dependencies. - `templates[].resource` applies, patches, or deletes Kubernetes resources. - `successCondition` and `failureCondition` define how Omnistrate decides whether a resource task completed. For Argo syntax details, see the [Argo Workflows documentation](https://argo-workflows.readthedocs.io/). Omnistrate renders `$sys.*`, `$var.*`, `$secret.*`, and `$func.*` expressions before executing the workflow against the tenant instance context. ## Optional Terraform Dependencies Operators often need cloud resources that are not Kubernetes resources, such as object-storage buckets, IAM roles, encryption keys, private networking, or DNS zones. Model those with a Terraform resource in the same service plan, mark it internal, and make the operator resource depend on it. The Terraform resource can export outputs, and the operator workflow can consume those outputs through normal Omnistrate parameters. This lets Terraform manage cloud primitives while the operator manages application lifecycle in Kubernetes. ## Build and Test Clone the template and build it with `omnistrate-ctl`: ``` git clone https://github.com/omnistrate-community/operator-spec-template cd operator-spec-template omnistrate-ctl build -f spec.yaml --name "Postgres Operator" --release-as-preferred --spec-type ServicePlanSpec ``` Then create a test instance from the Customer Portal or CLI. Validate: - The operator Helm dependencies are installed. - The tenant namespace is created. - The `create` workflow applies the operator custom resources. - Endpoints are visible after the operator reports readiness. - Manual backup creates a snapshot and stores provider metadata. - Restore creates a new target instance from the selected snapshot. - Start, stop, add capacity, and remove capacity drive the operator through its CRD. Use [Deployment Cell Access](https://docs.omnistrate.com/operate-guides/deployment-cell-access/index.md) and [Manage Workflows](https://docs.omnistrate.com/operate-guides/workflows/index.md) to debug Kubernetes resources and workflow events. For the broader architecture and integration principles, see [Beyond Operators: The Enterprise Control Plane Layer](https://omnistrate.com/blog/beyond-operators-the-enterprise-control-plane-layer) on the Omnistrate blog. # Build Guide Overview This build guide provides details on basic concepts. Omnistrate Build Guides are designed to help you configure, extend, and operate the private control plane generated for your SaaS product. Each guide covers a specific area of the build process, from initial setup and deployment models to managing Tenancy Types, Plans, Resources, and advanced configurations like dependencies, linking, and action hooks. - Deployment Models - Explore the supported deployment models (SaaS, PaaS, AaaS, BYOC, Air-Gapped) and how to configure them for your application. To learn more, see [here](https://docs.omnistrate.com/build-guides/deployment-models/index.md) - Tenancy Types - Understand different tenancy options (shared, siloed, hybrid) and how to select the right one for your customers. To learn more, see [here](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) - Plan - Define product plans with features, entitlements, and usage rules that drive customer subscriptions and billing. To learn more, see [here](https://docs.omnistrate.com/build-guides/plan/index.md) - Service Visibility - Control which services are exposed to which tenants, regions, or environments for fine‑grained visibility. To learn more, see [here](https://docs.omnistrate.com/build-guides/service-visibility/index.md) - Resource - Configure resources such as databases, compute, and storage that your application requires. To learn more, see [here](https://docs.omnistrate.com/build-guides/resource/index.md) - Dependencies - Set up inter‑service and inter‑resource dependencies to ensure correct startup, scaling, and failover sequences. To learn more, see [here](https://docs.omnistrate.com/build-guides/dependencies/index.md) - Resource Linking - Link resources together for integrated workflows, data flows, and cross‑service connectivity. To learn more, see [here](https://docs.omnistrate.com/build-guides/resource-linking/index.md) - Action Hooks - Use lifecycle hooks to run custom logic during provisioning, scaling, upgrades, or deprovisioning. To learn more, see [here](https://docs.omnistrate.com/build-guides/actionhooks/index.md) - Kubernetes Operators - Bring your operator-managed CRDs and model lifecycle operations with system and custom workflows. To learn more, see [here](https://docs.omnistrate.com/build-guides/operators/index.md) - API Parameters - Pass dynamic parameters to APIs and services for flexible runtime configuration. To learn more, see [here](https://docs.omnistrate.com/build-guides/api-params/index.md) - System Parameters - Leverage system‑generated parameters for consistent environment setup and automation. To learn more, see [here](https://docs.omnistrate.com/build-guides/system-parameters/index.md) Need Help? If you have questions, need clarification, or run into issues while following these guides, our team is here to help. Please reach out to [support@omnistrate.com](mailto:support@omnistrate.com) and we’ll be happy to assist. # Plan ## What is a Plan in Omnistrate A service consists of several Plans that allows you to offer different SaaS Product for different customer segments. ## How to use Plans in Omnistrate As part of the Plan, you can configure: - Deployment Models and Tenancy Types to meet different customer needs. For more information on Deployment Models and Tenancy Types, please see [here](https://docs.omnistrate.com/build-guides/deployment-models/index.md) and [here](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) - Plan details - Pricing - Support - Documentation - Subscription rules ### How it works After logging in to your Customer Portal, your users will be able to select the Plan and review the terms, documentation, support, pricing for the correspending Plan as configured by you. Once selected, they can then subscribe to the Plan and access the corresponding dashboard once approved (or auto-approved depending on how you configured your service). To learn more on subscription management, see [here](https://docs.omnistrate.com/tenant-management/subscription-management/index.md). ### Benefits Plan not only give options to your customers but also allows you to build custom GTM motion, pricing, support, feature set and cater to the needs of the respective customer segment. Naturally, as your customers will mature, they may need access to premium plans and can seamlessly migrate across different Plans allowing you to grow with their usage. ### Common Plans Most common Plans include: - Free tier offer - allows customers to try the product with limited features and capabilities - PAYG offer - allows customers to not have any commit but pay based on the usage - BYOC offer - allows customers to deploy application in their account - Dedicated or Enterprise offer - allows customers to have dedicated stack for extra security but at added cost. Usually, these plans have some base price or minimum commit associated with them - Serverless offer - allows customers to balance cost with security by having dedicated stack but not pay when not in use # Resource Linking ## What is Resource Linking An API parameter of resource type can be used to link any resource with another resource to enforce creating a linked resource before creating a parent resource. To understand, let's imagine if you have 2 resources and B is linked to A, then your users have to create A first, then B. In order to create B, they have to specify an instance of A that they want to link. You can also create a multi-level linkage and for your users to create a top-level resource, they have to follow the bottoms-up creation. Note To see the difference between resource linking and resource dependency, please see [here](#resource-linking-vs-resource-dependency) As an example, if you are creating a PostgreSQL SaaS such that you also want to allow proxy in front of their databases. Now, it makes no sense to create a proxy with no database to point to. To model this, you may create two Resources (or resources): - Database resource - Proxy resource And, then link Proxy resource with the Database resource. By doing that, Omnistrate will automatically link them and allow your users to specify a database instance before they can create a proxy instance. ## Resource Linking Example Here is an example with PGAdmin as a proxy resource and PostgreSQL as a database resource. ``` version: "3" services: Proxy: image: omnistrate/pgadmin4:7.5 ports: - 80:80 volumes: - ./data:/var/lib/pgadmin x-omnistrate-capabilities: httpReverseProxy: targetPort: 80 environment: - DB_ENDPOINT= Writer - SECURITY_CONTEXT_FS_GROUP=0 - SECURITY_CONTEXT_USER_ID=0 - SECURITY_CONTEXT_GROUP_ID=0 - PGADMIN_DEFAULT_EMAIL=test@gmail.com - PGADMIN_SERVER_JSON_FILE=/tmp/servers.json - PGADMIN_DEFAULT_PASSWORD=abc123 - DB_USERNAME=root x-omnistrate-api-params: - key: database description: backend database name: database type: Resource export: false required: true modifiable: false Database: image: 'bitnami/postgresql:latest' ports: - 5432:5432 volumes: - ./data:/var/lib/postgresql/data environment: - POSTGRESQL_PASSWORD=abc123 - POSTGRESQL_DATABASE=testdb - POSTGRESQL_USERNAME=root - POSTGRESQL_POSTGRES_PASSWORD=rootpassword12345 - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true - POSTGRESQL_REPLICATION_MODE=master - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=repl_password - POSTGRESQL_DATA_DIR=/var/lib/postgresql/data/dbdata - SECURITY_CONTEXT_USER_ID=1001 - SECURITY_CONTEXT_FS_GROUP=1001 - SECURITY_CONTEXT_GROUP_ID=0 ``` Assuming, you marked both the resources as external, your customers will be able to independently create an instance of database resource without any obligation to create a proxy resource, or configure any of the database instances to proxy instance. Note that your customers won't be able to create proxy instance without creating an instance of database resource because creation of proxy resource requires an instance of database instance. ## Resource Linking vs Resource Dependency Dependencies allow you to specify a DAG and how different Resources depend on each other. Resource linking is used to link two resources so that you can enforce creation of the linked resource before creating a parent resource. To better understand, let's imagine if you have 2 resources and B is linked to A, then your users have to create A first, then B. In order to create B, they have to specify an instance of A that they want to link. You can also create a multi-level linkage and for your users to create a top-level resource, they have to follow the bottoms-up creation by first creating the leaf resource in the DAG all the way up-to parent resource. In the dependency model, if you have B as a dependency to A, your users can directly create the parent resource A and Omnistrate will automatically detect the dependency, create B (aka all the dependencies) behind the scenes and use that to create A. Depending on the experience, you want to offer to your customers, you may want to choose one vs. the other. As an example, lets say you want to build a Database SaaS with Master and Replica. For linked resource, your customers will have to create Master instance and then add Replicas by passing the instance of Master instance during the creation of Replica instance(s). On the other hand, with dependency, you can offer a Cluster like experience and encapsulate both Master and Replica behind the Cluster resource. Now, your users can just the Cluster resource and Omnistrate will automatically create Master and Replica resources behind the scenes. There is no right or wrong and purely a function of what experience you want to offer. # Resource ## What is a Resource Resource (aka Resource) is a unit of functionality that can be run independently like a microservice. A Resource is comprised of several elements like image, infrastructure, configuration, capabilities etc. Let's go over each one of them in detail next. ## Elements of Resource - **Image** - allows you to configure image for this resource. Each resource can have a maximum of 1 image. - **Infrastructure** - allows you to configure infrastructure for this resource. Note that you can also parameterize your infrastructure based on the [API parameters](https://docs.omnistrate.com/build-guides/api-params/index.md). Each resource can have a maximum of 1 infrastructure. - **Environment variables** - Environment variables define a set of environment variables to configure your docker image. These environment variables can be static values or could be dynamic values based on the Input parameters. In addition, we have also defined some dynamic system parameters for your convenience. For more details, please see [this page](https://docs.omnistrate.com/build-guides/api-params/index.md) - **Configuration** - allows you to specify a configuration needed to run your docker image. Like environment variables, configuration can have static values or also dynamic values based on the Input parameters. For more details, please see [this page](https://docs.omnistrate.com/build-guides/api-params/index.md) - **Capabilities** - each component can be configured with set of capabilities as defined [here](https://docs.omnistrate.com/runtime-guides/overview/index.md) - **API parameters** - Your SaaS API parameters define a set of parameters that you want to take input from your customers. You can define the type, constraints, limits and we will make sure that those properties are honored in your generated control plane. For more details on API parameters, please see [this page](https://docs.omnistrate.com/build-guides/api-params/index.md) - **Action hooks** - these are custom logic that you may want to run at different stages of your Resource. As an example, if you want to install an extension to your PostgreSQL cluster, you can simply run an extension command on the post CLUSTER_INIT stage. Another example could be to run a rebalance command on adding a new node or removing a node. We will make sure that these commands are run reliably or alert you if we find any issues. For more details on action hooks, please see [this page](https://docs.omnistrate.com/build-guides/actionhooks/index.md) ## Types of Resource ### Passive vs Active If there is no image or infrastructure attached, we will treat that Resource as a metadata object to store information from your customers and configure other Resources. As an example, if you are building PostgreSQL SaaS and want to allow your users to input some of the PostgreSQL parameters, you can define a parameter-groups component to let your users specify list of parameters. We call such components as passive components. If image/infra are attached, we call such components as active components and automatically take care of the full lifecycle of such components as well. You can also combine different Resources to build more complex SaaS, ex - if you are building Kafka SaaS, you may have 1 Resource to represent Kafka brokers and another Resource to represent Zookeeper. Similarly, if you are building PostgreSQL HA SaaS, you may have 1 Resource to represent Master and another Resource to represent Replicas to configure Master and Replicas separately. ### Infrastructure type vs Deployment type Infrastructure type Resources support the underlying architecture and are crucial for the scalability and reliability of the service. These include Load Balancers, which can be of two types: Layer 4 (L4), which are TCP port-based and Layer 7 (L7), which are path-based and provide more advanced routing capabilities based on URL path. Additionally, file storage solutions are included as infrastructure resources, with current support limited to AWS. Deployment type Resources are the core building blocks of a SaaS provider's offering, integral to the data plane of the service. These components are essential for the operation and delivery of the service, handling the actual processing and management of user data. Additionally, serverless proxy Resources can be included in deployment resources to support serverless offerings. ### Internal vs External (tenant aware) A Resource can be internal or external. In your SaaS stack, there might be some components that you may want to directly expose to your customers while others may be technical details that you don’t want to expose to your customers. Omnistrate platform will build a SaaS experience such that only external (tenant aware) components will be exposed to your customers. In the above example, you can mark caching and database Resources as internal, and an application Resource as external. In this case, Omnistrate will only expose 1 Resource while generating the SaaS experience. Note Note that internal Resources doesn't imply that you can't open a port for your customers for those components. It just means that your control plane will not expose that resource to your customer to directly provision or other management operations. In fact, - You may have external ports open for your customers to interact. - Alternatively, you may have no ports open for your external service. ## Connecting Resources Your SaaS may have 1 or more Resources to form a multi-component application SaaS potentially [dependent on](https://docs.omnistrate.com/build-guides/dependencies/index.md) or [linked with](https://docs.omnistrate.com/build-guides/resource-linking/index.md) each other. ### Example Let’s say you want to build a LAMP stack based SaaS, you can define Memcached as your caching component, MySQL as a database Resource and an application server to host your PHP application as a third Resource. You can then connect them by marking caching and database Resources as dependencies to an application Resource. We will use this information to provision, scale, recover, patch different Resources accordingly. # Service Visibility ## What is Service Visibility Service visibility in Omnistrate allows you to control which services are internal and which are customer-facing. This configuration is essential when building multi-service applications. ## Service Visibility Configuration When defining multiple services in your compose file, you need to explicitly specify which services are internal using the `x-omnistrate-mode-internal` flag. This flag helps in managing service visibility and accessibility. ### Service Visibility Usage You can define service visibility in your compose file by adding the `x-omnistrate-mode-internal` flag to your service definition: ``` services: backend-service: x-omnistrate-mode-internal: true # This service will be internal # ...other service configuration public-api: x-omnistrate-mode-internal: false # This service will be customer-facing # ...other service configuration ``` Note When you have multiple services in your compose file, you must define the `x-omnistrate-mode-internal` flag for at least one service to explicitly specify its visibility. ### Configuration Rules - The flag accepts boolean values (`true`/`false`) - Required when defining multiple services in the compose file - At least one service must have this flag defined when multiple services are present - Helps in distinguishing between internal and customer-facing services ## Service Visibility Example Here's a complete example showing how to configure service visibility in a multi-service setup: ``` services: internal-backend: image: backend:latest x-omnistrate-mode-internal: true ports: - "8080" volumes: - internal-data:/data environment: - DB_CONNECTION=internal public-api-gateway: image: api-gateway:latest x-omnistrate-mode-internal: false ports: - "80:80" depends_on: - internal-backend ``` ## Service Visibility Best Practices - **Explicit Configuration**: Always explicitly define the visibility of your services to maintain clarity. - **Security Considerations**: - Mark internal services that handle sensitive operations as `x-omnistrate-mode-internal: true` - Only expose necessary services as customer-facing (`x-omnistrate-mode-internal: false`) - **Documentation**: Document the purpose and visibility of each service in your architecture ## Service Visibility Common Use Cases ### Internal Services Typically marked as `x-omnistrate-mode-internal: true`: - Database services - Cache services - Background workers - Internal APIs - Service meshes ### Customer-Facing Services Typically marked as `x-omnistrate-mode-internal: false`: - Public API endpoints - Customer-facing web applications - Load balancers - Gateway services # System Parameters ## What are System Parameters Omnistrate allows your deployments to be modified at runtime through contextual information provided by the platform through system parameters. System parameters are placeholders that are replaced with actual values at runtime. You can use system parameters in various places in your Helm charts, Kubernetes Operators, Container configurations, Kustomize templates, Terraform templates, or OpenTofu templates, etc. Example of information provided through system parameters (not exhaustive): - Cloud Provider - Region - Kubernetes Cluster ID - Tenant ID - VPC ID / Network ID - Subnet ID ## System Parameters Usage System parameters are used in the following format: `$sys.`. For example, `$sys.deploymentCell.region`. System parameters can be used when defining a service using Docker compose, Helm Charts (as part of the Helm Chart values), Operators (as part of the required Helm charts values or the operator configuration), Kustomize, Terraform templates or OpenTofu templates. Info All variables are substituted with their original data type. If you want to use a variable as a string, you need to wrap it in **escaped quotes**. For example, `\"$sys.deploymentCell.region\"`. This is particularly critical when using it on Helm Chart values that need a specific data type to be preserved / passed to the Helm Chart. ## Using Schema Validation for System Parameters You can use the following JSON schema in IDEs that use the YAML Language Server (eg: VSCode / NeoVim). ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/system-parameters-schema.json ``` For [**IntelliJ**](https://www.jetbrains.com/help/idea/yaml.html#use-schema-keyword) replace the top line with the following line to set up the yaml schema ``` # $schema: https://api.omnistrate.cloud/2022-09-01-00/schema/system-parameters-schema.json ``` This gives you an exhaustive list of all supported system parameters today. Please make sure to follow the syntax above to use it by prepending `$sys.` to the variable name. ## Expression Evaluator The [Expression Evaluator](https://docs.omnistrate.com/build-guides/evaluate-expressions/index.md) allows you to test and validate your expressions in real-time. You can use it to quickly prototype changes, debug issues, and ensure your expressions behave as expected before deploying them. ## Supported System Parameters The following table provide an extensive list of supported system parameters classified on - Compute: Information about the compute infrastructure on which the service instance is running - Network: Information about the network configuration and connectivity for each service instance - Storage: Information about storage configuration for the service instance - Tenant: Information about the tenant that requested the creation of the service - Deployment: Information about the deployment metadata, like Plan configuration - Deployment cell: Information about the runtime environment on which the service is hosted, like Kubernetes cluster information - Instance: Information about each particular service instance - Function: Miscellaneous utility functions to configure the service instance ### Compute parameters | **Parameter** | **Description** | **Type** | | ---------------------------------------- | ------------------------------------------------------------------- | -------- | | **`$sys.compute.node.poolName`** | Name of the node pool where the node is allocated. | string | | **`$sys.compute.node.cores`** | Number of CPU cores for the current node. | integer | | **`$sys.compute.node.memory`** | Amount of memory (RAM) in GB for the current node. | integer | | **`$sys.compute.node.instanceType`** | Instance type for the current node (e.g., `m5.large`). | string | | **`$sys.compute.node.name`** | Name of the current service node. | string | | **`$sys.compute.node.index`** | Index of the current node in the service instance. | integer | | **`$sys.compute.node.region`** | Cloud region where the current node is deployed. | string | | **`$sys.compute.node.version`** | Version of the node pool for the current node. | string | | **`$sys.compute.nodes[i].poolName`** | Name of the node pool where a specific node (node i) is allocated. | string | | **`$sys.compute.nodes[i].cores`** | Number of CPU cores for a specific node (node i). | integer | | **`$sys.compute.nodes[i].memory`** | Amount of memory (RAM) in GB for a specific node (node i). | integer | | **`$sys.compute.nodes[i].instanceType`** | Instance type for a specific node (node i) in the service instance. | string | | **`$sys.compute.nodes[i].name`** | Name of a specific node (node i) in the service instance. | string | | **`$sys.compute.nodes[i].index`** | Index of a specific node (node i) in the service instance. | integer | | **`$sys.compute.nodes[i].region`** | Cloud region where a specific node (node i) is deployed. | string | | **`$sys.compute.numNodes`** | Total number of nodes in the service instance. | integer | ### Network parameters | **Parameter** | **Description** | **Type** | | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | | **`$sys.network.node.internalEndpoint`** | Internal DNS endpoint name of the current node. | string | | **`$sys.network.node.externalEndpoint`** | External DNS endpoint name of the current node. | string | | **`$sys.network.node.availabilityZone.code`** | Code of the availability zone where the current node is located. | string | | **`$sys.network.node.availabilityZone.id`** | ID of the availability zone where the current node is located. | string | | **`$sys.network.node.internalIP`** | Internal IP address of the current node. | string | | **`$sys.network.node.hostIP`** | Host IP address of the current node. | string | | **`$sys.network.node.subnetID`** | Subnet ID associated with the current node. | string | | **`$sys.network.node.networkID`** | Network ID associated with the current node. | string | | **`$sys.network.node.cidrRange`** | CIDR range of the network for the current node. | string | | **`$sys.network.nodes[i].internalEndpoint`** | Internal network endpoint of a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].externalEndpoint`** | External network endpoint of a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].availabilityZone.code`** | Code of the availability zone for a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].availabilityZone.id`** | ID of the availability zone for a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].internalIP`** | Internal IP address of a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].hostIP`** | Host IP address of a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].subnetID`** | Subnet ID associated with a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].networkID`** | Network ID associated with a specific node (node i) in the network array. | string | | **`$sys.network.nodes[i].cidrRange`** | CIDR range of the network for a specific node (node i) in the network array. | string | | **`$sys.network.internalClusterEndpoint`** | Internal endpoint of the cluster. | string | | **`$sys.network.internalClusterServerlessEndpoint.endpointName`** | Endpoint name for the internal cluster serverless endpoint. | string | | **`$sys.network.internalClusterServerlessEndpoint.openPorts[i]`** | Open ports for the internal cluster serverless endpoint (port 0). | integer | | **`$sys.network.internalClusterServerlessEndpoint.partitionID`** | Partition ID for the internal cluster serverless endpoint. | string | | **`$sys.network.externalClusterEndpoint`** | External endpoint of the cluster. | string | | **`$sys.network.externalClusterServerlessEndpoint.endpointName`** | Endpoint name for the external cluster serverless endpoint. | string | | **`$sys.network.externalClusterServerlessEndpoint.openPorts[i]`** | Open ports for the external cluster serverless endpoint (port 0). | integer | | **`$sys.network.externalClusterServerlessEndpoint.partitionID`** | Partition ID for the external cluster serverless endpoint. | string | | **`$sys.network.availabilityZones[i].code`** | Code of the availability zone in the array of availability zones. | string | | **`$sys.network.availabilityZones[i].id`** | ID of the availability zone in the array of availability zones. | string | | **`$sys.network.customDNS`** | Resolves to the Resource's configured custom DNS when the `customDNS` capability is enabled. If no custom DNS is configured, it falls back to `$sys.network.externalClusterEndpoint`. Prefer this for application-facing hostnames such as CORS allow-lists, callback URLs, and `ALLOWED_HOSTS`. | string | | **`$sys.network.namedPorts`** | Map of named ports to port/port-ranges configured for the cluster. Access individual named ports using bracket notation, for example `$sys.network.namedPorts["port-name"]`. | string | #### Using `$sys.network.customDNS` When your application needs to know the public hostname that customers will use, prefer `$sys.network.customDNS` over node-level endpoints such as `$sys.network.node.externalEndpoint`. Common examples include: - `ALLOWED_HOSTS` - CORS allow-lists - OAuth callback URLs - Frontend base URLs Behavior notes: - Enable `x-omnistrate-capabilities.customDNS` on the Resource to allow custom DNS aliases. - If no custom DNS is configured for the instance, `$sys.network.customDNS` automatically falls back to `$sys.network.externalClusterEndpoint`. - If you add or change custom DNS after the instance is already deployed, restart the deployment instance so environment variables and rendered configuration are refreshed. ### Storage parameters | **Parameter** | **Description** | **Type** | | --------------------------------------- | ------------------------------------------------------------- | -------- | | **`$sys.storage.volumes[i].name`** | Name of the storage volume (volume i) in the volumes array. | string | | **`$sys.storage.volumes[i].size`** | Size of the storage volume (volume i) in GB. | integer | | **`$sys.storage.volumes[i].type`** | Type of the storage volume (volume i) (e.g., SSD, HDD). | string | | **`$sys.storage.volumes[i].mountPath`** | Mount path for the storage volume (volume i). | string | | **`$sys.storage.volumes[i].id`** | ID of the storage volume (volume i). | string | | **`$sys.storage.volumes[i].pvName`** | Persistent Volume (PV) name of the storage volume (volume i). | string | | **`$sys.storage.numVolumes`** | Total number of storage volumes in the current node. | integer | ### Tenant parameters | **Parameter** | **Description** | **Type** | | ------------------------- | --------------------------------------------- | -------- | | **`$sys.tenant.userID`** | User ID of the tenant. | string | | **`$sys.tenant.name`** | Name of the tenant. | string | | **`$sys.tenant.email`** | Email address of the tenant. | string | | **`$sys.tenant.orgId`** | Organization ID associated with the tenant. | string | | **`$sys.tenant.orgName`** | Organization name associated with the tenant. | string | ### Deployment parameters | **Parameter** | **Description** | **Type** | | ---------------------------------------------------- | --------------------------------------------------------------------------- | -------- | | **`$sys.deployment.serviceID`** | ID of the service. | string | | **`$sys.deployment.planID`** | ID of the deployment plan. | string | | **`$sys.deployment.planName`** | Name of the deployment plan. | string | | **`$sys.deployment.planVersion`** | Version of the deployment plan. | string | | **`$sys.deployment.resourceAlias`** | Alias for the deployed resource. | string | | **`$sys.deployment.resourceID`** | ID of the deployed resource. | string | | **`$sys.deployment.resourceKubernetesNamespace`** | Kubernetes namespace where the deployed resource is created. | string | | **`$sys.deployment.nodeAffinityRules`** | Node affinity rules for the deployment. | string | | **`$sys.deployment.nodeSelectorRules`** | Node selector rules for the deployment. | string | | **`$sys.deployment.tlsServerCertificateSecretName`** | Secret name for the TLS server certificate. | string | | **`$sys.deployment.tlsCertMountPath`** | Mount path for the TLS certificate in the deployment. | string | | **`$sys.deployment.imageNameWithTag`** | Full image name with the tag for the deployment container. | string | | **`$sys.deployment.imagePullSecretName`** | Name of the image pull secret used for the deployment container image. | string | | **`$sys.deployment.iamWorkloadRoleARN`** | ARN of the IAM workload role for the deployment. | string | | **`$sys.deployment.kubernetesServiceAccountName`** | Name of the Kubernetes service account used in the deployment. | string | | **`$sys.deployment.cloudProvider`** | Cloud provider where the deployment is hosted (e.g., AWS, GCP, Azure, OCI). | string | | **`$sys.deployment.cloudProviderAccountID`** | Account ID of the cloud provider. | string | | **`$sys.deployment.gcpProjectNumber`** | Project number for the deployment if using GCP. | string | | **`$sys.deployment.gcpBootstrapEmail`** | Email used for GCP bootstrap actions in the deployment. | string | | **`$sys.deployment.kubernetesClusterID`** | ID of the Kubernetes cluster used for the deployment. | string | | **`$sys.deployment.environmentID`** | ID of the environment used for the deployment. | string | | **`$sys.deployment.environmentType`** | Type of the environment used for the deployment (e.g., Prod, Dev). | string | | **`$sys.deployment.azureTenantID`** | Azure Tenant ID (Entra identifier) used for the deployment. | string | | **`$sys.deployment.ociDomainID`** | OCI Domain ID used for the deployment. | string | ### Deployment cell parameters | **Parameter** | **Description** | **Type** | | ------------------------------------------------- | -------------------------------------------------------------------------------- | -------- | | **`$sys.deploymentCell.cloudProviderName`** | Name of the cloud provider for the deployment cell (e.g., AWS, GCP, Azure, OCI). | string | | **`$sys.deploymentCell.region`** | Region where the deployment cell is located. | string | | **`$sys.deploymentCell.kubernetesClusterID`** | The ID of the Kubernetes cluster corresponding to the deployment cell. | string | | **`$sys.deploymentCell.cloudProviderNetworkID`** | Network ID of the cloud provider for the deployment cell. | string | | **`$sys.deploymentCell.publicSubnetIDs[i].id`** | ID of the public subnet in the deployment cell. | string | | **`$sys.deploymentCell.privateSubnetIDs[i].id`** | ID of the private subnet in the deployment cell. | string | | **`$sys.deploymentCell.availabilityZones[i].id`** | ID of the availability zone in the deployment cell. | string | | **`$sys.deploymentCell.subNetworks`** | List of sub network IDs in the deployment cell. | string | | **`$sys.deploymentCell.securityGroupID`** | Security group ID for the deployment cell. | string | | **`$sys.deploymentCell.oidcIssuerID`** | OIDC Issuer ID for the deployment cell. | string | | **`$sys.deploymentCell.cloudProviderAccountID`** | Account ID of the cloud provider of the deployment cell. | string | | **`$sys.deploymentCell.environmentType`** | Type of the environment that the deployment cell is used for. | string | #### Cross-cloud notes for deployment cell networking - `$sys.deploymentCell.cloudProviderNetworkID` is the provider-native network identifier, such as a VPC or VNet. - `$sys.deploymentCell.publicSubnetIDs[*].id` and `$sys.deploymentCell.privateSubnetIDs[*].id` are the usual subnet fields to use on AWS and Azure. - On GCP, prefer `$sys.deploymentCell.subNetworks` instead of assuming `publicSubnetIDs` or `privateSubnetIDs` are populated. - `$sys.deploymentCell.kubernetesClusterID` is the provider cluster identifier and can differ from the Omnistrate deployment cell ID. - For Terraform and other pre-Kubernetes resources, prefer `$sys.deploymentCell.*` rather than node-scoped compute or network fields. - If you are not sure which fields are populated in a live resource context, use the [Expression Evaluator](https://docs.omnistrate.com/build-guides/evaluate-expressions/index.md). #### AWS specific deployment cell parameters | **Parameter** | **Description** | **Type** | | -------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | **`$sys.deploymentCell.aws.accountNumber`** | Account number of the AWS account hosting deployment cell (this value will be the same as `$sys.deploymentCell.cloudProviderAccountID`). | string | | **`$sys.deploymentCell.aws.certManagerIAMRoleARN`** | ARN of the IAM role used by cert-manager in the deployment cell. | string | | **`$sys.deploymentCell.aws.managedWorkloadIdentities["managedIdentityIdentifier"].roleARN`** | ARN of the IAM role for the AWS managed workload identity (if used for account). See [Managed Workload Identities](https://docs.omnistrate.com/operate-guides/managed-workload-identities/index.md). | string | #### GCP specific deployment cell parameters | **Parameter** | **Description** | **Type** | | -------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | **`$sys.deploymentCell.gcp.projectID`** | Project ID of the GCP project hosting deployment cell (this value will be the same as `$sys.deploymentCell.cloudProviderAccountID`). | string | | **`$sys.deploymentCell.gcp.projectNumber`** | Project number of the GCP project hosting deployment cell. | string | | **`$sys.deploymentCell.gcp.bootstrapEmail`** | Email used for bootstrap service principal responsible for managing deployment cell. | string | | **`$sys.deploymentCell.gcp.certManagerServiceAccountEmail`** | Service account email used by cert-manager in the deployment cell. | string | | **`$sys.deploymentCell.gcp.managedWorkloadIdentities["managedIdentityIdentifier"].serviceAccountEmail`** | GCP service account email for the managed workload identity (if configured for the account). See [Managed Workload Identities](https://docs.omnistrate.com/operate-guides/managed-workload-identities/index.md). | string | #### Azure specific deployment cell parameters | **Parameter** | **Description** | **Type** | | ----------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | **`$sys.deploymentCell.azure.subscriptionID`** | Subscription ID of the Azure subscription hosting deployment cell (this value will be the same as `$sys.deploymentCell.cloudProviderAccountID`). | string | | **`$sys.deploymentCell.azure.tenantID`** | Tenant ID (Entra identifier) used by the deployment cell. | string | | **`$sys.deploymentCell.azure.terraformClientID`** | Client ID of the service principal used for terraform resources. | string | | **`$sys.deploymentCell.azure.bootstrapClientID`** | Client ID of the service principal used to manage deployment cell. | string | | **`$sys.deploymentCell.azure.dataplaneAgentClientID`** | Client ID of the service principal used to operate deployment cell. | string | | **`$sys.deploymentCell.azure.dataplaneNodeClientID`** | Client ID of the service principal used as identity for dataplane nodes on the deployment cell. | string | | **`$sys.deploymentCell.azure.certManagerClientID`** | Client ID used by cert-manager in the deployment cell. | string | | **`$sys.deploymentCell.azure.managedWorkloadIdentities["managedIdentityIdentifier"].clientID`** | Application client ID for the Azure managed workload identity (if used for account). See [Managed Workload Identities](https://docs.omnistrate.com/operate-guides/managed-workload-identities/index.md). | string | | **`$sys.deploymentCell.azure.managedWorkloadIdentities["managedIdentityIdentifier"].tenantID`** | Entra tenant ID for the Azure managed workload identity. See [Managed Workload Identities](https://docs.omnistrate.com/operate-guides/managed-workload-identities/index.md). | string | #### OCI specific deployment cell parameters | **Parameter** | **Description** | **Type** | | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | **`$sys.deploymentCell.oci.tenancyOCID`** | Tenancy OCID of the deployment cell (this value will be the same as `$sys.deploymentCell.cloudProviderAccountID`). | string | | **`$sys.deploymentCell.oci.domainOCID`** | Domain OCID of the deployment cell. | string | | **`$sys.deploymentCell.oci.domainHomeRegion`** | Home region of the OCI identity domain used by the deployment cell account. | string | | **`$sys.deploymentCell.oci.accountCompartmentOCID`** | Compartment OCID of the compartment created during account onboarding. | string | | **`$sys.deploymentCell.oci.clusterCompartmentOCID`** | Compartment OCID created for deployment-cell resources. In the default hierarchy it is a child of `$sys.deploymentCell.oci.accountCompartmentOCID`; with separate network stacks it is created under dedicated network compartment. | string | ### Instance parameters | **Parameter** | **Description** | **Type** | | ------------- | ------------------------------------------- | -------- | | **`$sys.id`** | Unique identifier for the service instance. | string | ### Functions | **Function** | **Parameters** | **Description** | **Return Type** | **Example Usage** | | -------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `$func.uuidv4` | none | Generates a UUID v4 | string | {{ `$func.uuidv4()` }} → `"550e8400-e29b-41d4-a716-446655440000"` | | `$func.random` | `type` (string | int64), `length/max` (number), `seed` (optional number) | Generates random value. For string type, generates random string of given length. For int64 type, generates random number up to max. Optional seed for deterministic results | string | | `$func.randomminmax` | `min` (number), `max` (number), `seed` (optional number) | Generates random number between min and max. Optional seed for deterministic results | string | {{ `$func.randomminmax(1000, 2000, 1234)`}} → `"1682"` | | `$func.randomport` | `port name` (from namedPorts map), `seed` (optional number) | Generates random port from named port specification. Optional seed for deterministic results | string | "{{ `$func.randomport($sys.network.namedPorts['grpc-range'], 1234)` }}" → `"1682"` | | `$func.sha256` | `data` (string), `salt` (optional string) | Generates SHA256 hash of input data. If salt provided, hashes "salt.data" | string (64 chars hex) | {{ `$func.sha256("hello")` }} → `"2cf24dba..."` {{ `$func.sha256("hello", "salt")` }} → `"4c5fd1b5..."` | | `$func.add` | `num1` (number), `num2` (number) | Adds two numbers together | string | {{ `$func.add(5, 3)` }} → `"8"` | | `$func.toupper` | `str` (string) | Converts string to uppercase | string | {{ `$func.toupper("hello")` }} → `"HELLO"` | | `$func.tolower` | `str` (string) | Converts string to lowercase | string | {{ `$func.tolower("HELLO")` }} → `"hello"` | | `$func.length` | `str` (string) | Returns length of input string | string | {{ `$func.length("hello")` }} → `"5"` | | `$func.max` | `...numbers` (variadic numbers) | Returns maximum value from provided numbers | string | {{ `$func.max(1, 5, 3)` }} → `"5"` | | `$func.min` | `...numbers` (variadic numbers) | Returns minimum value from provided numbers | string | {{ `$func.min(1, 5, 3)` }} → `"1"` | | `$func.concat` | `str1` (string), `str2` (string) | Concatenates two strings together | string | {{ `$func.concat("hello", "world")` }} → `"helloworld"` | | `$func.substring` | `str` (string), `start` (number), `end` (number) | Returns substring from start index to end index | string | {{ `$func.substring("hello", 1, 4)` }} → `"ell"` | | `$func.truncate` | `str` (string), `length` (number) | Truncates string to specified length | string | {{ `$func.truncate("hello", 3)` }} → `"hel"` | | `$func.foreach` | `array` (array), `template` (string) | Applies template to each array element and joins results with commas | string | {{ `$func.foreach(["a","b","c"], "{{.}}.test")` }} → `"a.test,b.test,c.test"` | | `$func.mult` | `number1 (float64), number2 (float64)` | Multiplies two numbers together | string (number) | {{ `$func.mult(5, 7)` }} → `"35"` | | `$func.mult` | `number1 (float64), number2 (float64)` | Multiplies two numbers together | string (number) | {{ `$func.mult(5.5, 7)` }} → `"38.5"` | | `$func.div` | `number1 (float64), number2 (float64)` | Divides the first number by the second number. Returns "NaN" if dividing by zero. | string (number or "NaN") | {{ `$func.div(10, 2)` }} → `"5"` | | `$func.div` | `number1 (float64), number2 (float64)` | Divides the first number by the second number. Returns "NaN" if dividing by zero. | string (number or "NaN") | {{ `$func.div(9, 2)` }} → `"4.5"` | | `$func.base64encode` | `data (string)` | Encodes a string to base64 | string | {{ `$func.base64encode("hello world")` }} → `"aGVsbG8gd29ybGQ="` | | `$func.base64decode` | `data (string)` | Decodes a base64 string | string | {{ `$func.base64decode("aGVsbG8gd29ybGQ=")` }} → `"hello world"` | | `$func.equals` | `value1` (any), `value2` (any) | Compares two values for equality. Quotes around literal arguments are ignored, so `"aws"` and `aws` compare equal | string ("true" or "false") | {{ `$func.equals($sys.deploymentCell.cloudProviderName, "aws")` }} → `"true"` on AWS deployment cells | | `$func.if` | `condition` (boolean), `valueIfTrue` (string), `valueIfFalse` (string) | Returns `valueIfTrue` when the condition evaluates to `true`, `valueIfFalse` otherwise. Combine with `$func.equals` to express conditional values | string | {{ `$func.if($func.equals($var.deploymentMode, "replicated"), "3", "1")` }} → `"3"` when `$var.deploymentMode` is `replicated`, `"1"` otherwise | # Tenancy Types ## What Tenancy Types are Supported by Omnistrate Omnistrate supports three distinct tenancy types to meet different isolation, cost, and operational requirements: - **Dedicated Tenancy** (`OMNISTRATE_DEDICATED_TENANCY`): When using this tenancy type in your Docker Compose specification, each tenant deployment instance gets dedicated infrastructure - **Multi Tenancy** (`OMNISTRATE_MULTI_TENANCY`): When using this type in your Docker Compose specification, multiple deployment instances, from different tenants, share infrastructure resources with logical isolation - **Custom Tenancy** (`CUSTOM_TENANCY`): For Helm, Kustomize, Terraform and OpenTofu deployments where tenancy is defined within the components themselves and not controlled by Omnistrate Each tenancy type serves different use cases and provides varying levels of isolation, cost efficiency, and operational complexity. These tenancy types describe Hosted SaaS deployments. For the available Hosted SaaS tenancy options, see the [Cellular Multi-Tenancy Overview](https://docs.omnistrate.com/build-guides/cellular-multi-tenancy-overview/index.md) and [Dedicated Tenancy Overview](https://docs.omnistrate.com/build-guides/dedicated-tenancy-overview/index.md). ## Dedicated Tenancy In Dedicated tenancy model, each tenant deployment receives their own dedicated infrastructure stack (VMs). Omnistrate manages the provisioning, scaling, and lifecycle of these dedicated resources while ensuring complete isolation between tenants. ### Dedicated cells You can go even further by configuring `maximumDeploymentsPerCell` to ensure a dedicated host cluster for each deployment. To learn more, see [Custom Deployment Cell Placement](https://docs.omnistrate.com/infra-guides/custom-deployment-cell-placement/index.md). ### Dedicated Network In addition, you can also Custom Network to ensure complete isolation from network, deployment cell, VM machines and attached storage. You can read more [here](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md). ### How Omnistrate Dedicated Tenancy Works **Complete Isolation**: Each tenant gets their own virtual machines, storage, and network resources. No sharing occurs between tenants at the infrastructure level. **Customization**: Tenants can customize their infrastructure configuration, including instance types, storage options, and network settings. **Security and Compliance**: Provides the highest level of security and is suitable for enterprise customers with strict compliance requirements. **Omnistrate Management**: While infrastructure is dedicated per tenant, Omnistrate still handles provisioning, monitoring, scaling, and maintenance operations. **Use Cases**: Dedicated tenancy is ideal for enterprise customers, regulated industries, high-security requirements, and scenarios where performance isolation is critical. Many SaaS products like Amazon RDS, Confluent Cloud, and MongoDB Atlas dedicated clusters use this model. **Configuration Example**: ``` x-omnistrate-service-plan: tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' ``` ## Multi-Tenancy In Omnistrate's multi-tenancy model, multiple customer instances share the same underlying infrastructure while maintaining logical isolation. Omnistrate handles the orchestration, resource allocation, and tenant isolation automatically. ### How Omnistrate Multi-Tenancy Works **Resource Sharing and Bin-Packing**: Omnistrate automatically places multiple tenant instances on the same virtual machines based on resource requirements (CPU, memory). This bin-packing approach maximizes resource utilization and reduces costs. **Automatic Scaling**: When resource demands exceed the capacity of existing VMs, Omnistrate automatically provisions new virtual machines and redistributes workloads as needed. **Logical Isolation**: Each tenant's workload runs in isolated containers with defined resource limits and requests, ensuring one tenant cannot impact another's performance. **Configuration Example**: ``` x-omnistrate-service-plan: tenancyType: 'OMNISTRATE_MULTI_TENANCY' services: postgres: deploy: resources: limits: cpus: '0.50' memory: 50M reservations: cpus: '0.25' memory: 20M ``` You can also configure resource requests and limits as API parameters: ``` services: postgres: x-omnistrate-compute: resourceRequestMemoryAPIParam: resourceRequestMemory resourceRequestCPUAPIParam: resourceRequestCPU resourceLimitMemoryAPIParam: resourceLimitMemory resourceLimitCPUAPIParam: resourceLimitCPU x-omnistrate-api-params: - key: resourceRequestMemory name: Memory Request description: Memory Request for the Postgres instance type: String modifiable: true required: true export: true defaultValue: 100M - key: resourceRequestCPU name: CPU Request description: CPU Request for the Postgres instance type: String modifiable: true required: true export: true defaultValue: 100m - key: resourceLimitMemory name: Memory Limit description: Memory Limit for the Postgres instance type: String modifiable: true required: true export: true defaultValue: 2048M - key: resourceLimitCPU name: CPU Limit description: CPU Limit for the Postgres instance type: String modifiable: true required: true export: true defaultValue: "4" ``` **Architecture Selection**: You can specify processor architecture using the `platform` attribute: ``` services: postgres: platform: "linux/arm64" # Default is X86_64 ``` **Use Cases**: Multi-tenancy is ideal for cost-effective SaaS offerings, free tiers, and scenarios where moderate isolation is sufficient. Many SaaS products like Redis Cloud or MongoDB Atlas free tiers use this model. ### In-App Multi-Tenancy If your application is already tenant-aware, you can group tenants into cells for enhanced security and fault tolerance. Each cell is a single application instance supporting a small group of tenants. Many SaaS products like HubSpot, Pulley, and Zendesk use this cell-based multi-tenancy model. ## Custom Tenancy Custom tenancy is used for Helm, Kustomize, Terraform and OpenTofu deployments where the tenancy model is defined within the deployment components themselves, rather than being controlled by Omnistrate. ### How Custom Tenancy Works **Custom-Defined Tenancy**: The tenancy model, isolation boundaries, and resource allocation are defined within your Helm charts, Terraform configurations, or Kustomize overlays. **Omnistrate Orchestration**: Omnistrate handles the deployment, lifecycle management, and scaling of your components, but respects the tenancy rules you've defined. **Flexibility**: You have complete control over how tenants are isolated, how resources are shared or dedicated, and how scaling occurs. **Component-Level Control**: Different components within your service can have different tenancy models based on your specific requirements. **Use Cases**: Custom tenancy is ideal when you have specific tenancy requirements that don't fit the standard multi-tenant or dedicated models, need fine-grained control over resource allocation, or are migrating existing infrastructure-as-code deployments to Omnistrate. ## Choosing the Right Tenancy Type | Tenancy Type | Isolation Level | Cost Efficiency | Management Overhead | Best For | | ------------------------------ | --------------- | --------------- | ------------------- | -------------------------------------------------------------------- | | `OMNISTRATE_DEDICATED_TENANCY` | Infrastructure | Low | Low | Enterprise customers, compliance requirements, performance isolation | | `OMNISTRATE_MULTI_TENANCY` | Logical | High | Low | Cost-effective SaaS, free tiers, moderate isolation needs | | `CUSTOM_TENANCY` | User-defined | Variable | Medium | Specific tenancy requirements, existing IaC, fine-grained control | The choice of tenancy type depends on your specific requirements for isolation, cost, compliance, and operational complexity. You can also mix different tenancy types within the same service for different tiers or customer segments. # Multi-Cloud Terraform Configuration ## Overview Omnistrate allows you to define separate Terraform stacks for each cloud provider — AWS, GCP, Azure, OCI, and Nebius — within a single Plan specification. At deployment time, the platform automatically selects and executes the correct stack based on the target cloud provider, letting you offer a multi-cloud SaaS Product from one specification. This approach avoids complex conditionals inside a single Terraform configuration. Instead, you maintain clean, provider-specific stacks that follow each cloud's best practices. ## Configuring Per-Cloud-Provider Stacks The `configurationPerCloudProvider` property under `terraformConfigurations` accepts entries for `aws`, `gcp`, `azure`, `oci`, and `nebius`. Each entry points to a Terraform stack in your Git repository. ### Basic Structure ``` services: - name: cloudInfra internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git gcp: terraformPath: /terraform/gcp gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git azure: terraformPath: /terraform/azure gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git oci: terraformPath: /terraform/oci gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git nebius: terraformPath: /terraform/nebius serviceAccountID: serviceaccount-e00vqdp9fskhmmaan8 publicKeyID: publickey-e00h9scsyy9mbefrjf privateKeyPEM: $secret.nebiusTerraformPrivateKey gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Each cloud provider entry requires: - **`terraformPath`**: The directory path within the repository containing the Terraform stack for that provider - **`gitConfiguration`**: The Git repository URL and branch/tag reference ### Repository Layout A common repository structure for multi-cloud Terraform stacks: ``` infra-repo/ ├── terraform/ │ ├── aws/ │ │ ├── main.tf │ │ ├── variables.tf │ │ └── outputs.tf │ ├── gcp/ │ │ ├── main.tf │ │ ├── variables.tf │ │ └── outputs.tf │ ├── oci/ │ │ ├── main.tf │ │ ├── variables.tf │ │ └── outputs.tf │ ├── nebius/ │ │ ├── main.tf │ │ ├── variables.tf │ │ └── outputs.tf │ └── azure/ │ ├── main.tf │ ├── variables.tf │ └── outputs.tf ``` Tip Keep consistent output names across cloud providers. If your AWS stack outputs `database_endpoint`, your GCP, Azure, OCI, and Nebius stacks should output the same key. This ensures dependent resources can reference outputs without cloud-specific logic. ## Cloud-Specific Examples ### AWS Stack ``` provider "aws" { region = "{{ $sys.deploymentCell.region }}" } resource "aws_db_instance" "database" { identifier = "db-{{ $sys.id }}" engine = "postgres" instance_class = "db.t3.medium" allocated_storage = 20 db_subnet_group_name = aws_db_subnet_group.main.name vpc_security_group_ids = [aws_security_group.db.id] username = "admin" password = var.db_password skip_final_snapshot = true } resource "aws_security_group" "db" { name = "db-sg-{{ $sys.id }}" vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" ingress { from_port = 5432 to_port = 5432 protocol = "tcp" cidr_blocks = ["{{ $sys.deploymentCell.cidrRange }}"] } } resource "aws_db_subnet_group" "main" { name = "db-subnet-{{ $sys.id }}" subnet_ids = [ "{{ $sys.deploymentCell.privateSubnetIDs[0].id }}", "{{ $sys.deploymentCell.privateSubnetIDs[1].id }}" ] } output "database_endpoint" { value = aws_db_instance.database.endpoint } output "database_port" { value = aws_db_instance.database.port } ``` ### GCP Stack ``` provider "google" { region = "{{ $sys.deploymentCell.region }}" } resource "google_sql_database_instance" "database" { name = "db-{{ $sys.id }}" database_version = "POSTGRES_15" region = "{{ $sys.deploymentCell.region }}" settings { tier = "db-f1-micro" ip_configuration { ipv4_enabled = false private_network = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" } } deletion_protection = false } resource "google_sql_user" "admin" { name = "admin" instance = google_sql_database_instance.database.name password = var.db_password } output "database_endpoint" { value = google_sql_database_instance.database.private_ip_address } output "database_port" { value = 5432 } ``` ### Azure Stack ``` provider "azurerm" { features {} } resource "azurerm_postgresql_flexible_server" "database" { name = "db-{{ $sys.id }}" resource_group_name = var.resource_group_name location = "{{ $sys.deploymentCell.region }}" version = "15" administrator_login = "admin" administrator_password = var.db_password sku_name = "B_Standard_B1ms" storage_mb = 32768 zone = "1" } output "database_endpoint" { value = azurerm_postgresql_flexible_server.database.fqdn } output "database_port" { value = 5432 } ``` ### Nebius Provider Block Nebius Terraform stacks should declare the Nebius provider, but leave credential injection to Omnistrate: ``` terraform { required_providers { nebius = { source = "terraform-provider.storage.eu-north1.nebius.cloud/nebius/nebius" } } } provider "nebius" { domain = "api.eu.nebius.cloud:443" } ``` Configure the credentials for that stack on the Plan entry by setting `serviceAccountID`, `publicKeyID`, and `privateKeyPEM` under `configurationPerCloudProvider.nebius`. ### OCI Stack ``` provider "oci" { region = "{{ $sys.deploymentCell.region }}" } resource "oci_database_autonomous_database" "database" { compartment_id = "{{ $sys.deploymentCell.oci.clusterCompartmentOCID }}" display_name = "db-{{ $sys.id }}" db_name = "db{{ $sys.id }}" admin_password = var.db_password cpu_core_count = 1 data_storage_size_in_tbs = 1 db_version = "19c" db_workload = "OLTP" is_free_tier = false } output "database_endpoint" { value = oci_database_autonomous_database.database.connection_urls[0].sql_dev_web_url } output "database_port" { value = 1522 } ``` ## Using System Parameters Across Clouds Omnistrate provides cloud-agnostic system parameters that work across all providers. Use these to inject deployment context into your Terraform templates: | Parameter | Description | | -------------------------------------------- | ------------------------------------------------------------------ | | `$sys.deploymentCell.region` | The deployment region (e.g., `us-east-1`, `us-central1`, `eastus`) | | `$sys.deploymentCell.cloudProviderNetworkID` | The VPC/VNet/Network ID for the deployment cell | | `$sys.deploymentCell.cidrRange` | The CIDR range assigned to the deployment cell | | `$sys.deploymentCell.publicSubnetIDs[i].id` | Public subnet IDs available in the deployment cell | | `$sys.deploymentCell.privateSubnetIDs[i].id` | Private subnet IDs available in the deployment cell | | `$sys.id` | Unique deployment identifier, useful for naming resources | For the full list, see [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). Warning Ensure the main `.tf` file for each cloud provider stack includes the provider definition. Omnistrate requires the provider block to be present in the top-level Terraform file. ## Versioning with Git Tags Use Git tags to version your Terraform stacks and ensure consistency across Plan releases: ``` terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/tags/v1.2.0 repositoryUrl: https://github.com/your-org/infra-repo.git ``` Each Plan version can reference a specific Git tag, making deployments reproducible, and upgrades controlled. ## Private Repository Access For private Git repositories, provide an access token: ``` gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/private-infra-repo.git accessToken: ``` ## Custom Execution Identity For Hosted SaaS deployments, Omnistrate manages Terraform execution identities for each cloud provider. For AWS, GCP, Azure, and OCI, the identity is auto-created — you only need to assign the permissions your Terraform stack requires. You do not need to specify the identity in the Plan spec. - **AWS** (auto-created): Omnistrate creates an IAM role named `omnistrate-terraform-execution-role` in your AWS account. Assign the permissions your Terraform stack needs to this role. - **GCP** (auto-created): Omnistrate creates a service account named `omnistrate-tf-` in your GCP project. Assign the permissions your Terraform stack requires to this service account. - **Azure** (auto-created): Omnistrate creates a service principal named `terraform--`. Assign the permissions your Terraform stack requires to this service principal. - **OCI** (auto-created): Omnistrate creates a user named `-terraform-user`. Assign the permissions your Terraform stack requires to this user. ``` terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git gcp: terraformPath: /terraform/gcp gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Nebius is different. Use explicit service-account auth on the `nebius` entry: ``` terraformConfigurations: configurationPerCloudProvider: nebius: terraformPath: /terraform/nebius serviceAccountID: "serviceaccount-e00vqdp9fskhmmaan8" publicKeyID: "publickey-e00h9scsyy9mbefrjf" privateKeyPEM: "$secret.nebiusTerraformPrivateKey" gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Notes: - All three Nebius auth fields are required on `configurationPerCloudProvider.nebius`. - Prefer `$secret.` for `privateKeyPEM`. For more details on custom execution identities, pre-created principals, and BYOC Terraform permissions, see the [Getting Started with Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/#additional-permissions-for-terraform) guide. ## Full Multi-Cloud Example The following Plan specification defines a Terraform resource that provisions a database on each cloud, with a Helm chart application consuming the database endpoint: ``` name: Multi-Cloud SaaS Product deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" services: - name: dbInfra internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/tags/v1.0.0 repositoryUrl: https://github.com/your-org/infra-repo.git gcp: terraformPath: /terraform/gcp gitConfiguration: reference: refs/tags/v1.0.0 repositoryUrl: https://github.com/your-org/infra-repo.git nebius: terraformPath: /terraform/nebius serviceAccountID: "serviceaccount-e00vqdp9fskhmmaan8" publicKeyID: "publickey-e00h9scsyy9mbefrjf" privateKeyPEM: "$secret.nebiusTerraformPrivateKey" gitConfiguration: reference: refs/tags/v1.0.0 repositoryUrl: https://github.com/your-org/infra-repo.git - name: WebApp dependsOn: - dbInfra network: ports: - 8080 helmChartConfiguration: chartName: web-app chartVersion: 2.0.0 chartRepoName: my-charts chartRepoURL: https://charts.example.com chartValues: database: host: "{{ $dbInfra.out.database_endpoint }}" port: "{{ $dbInfra.out.database_port }}" ``` When a customer deploys this Plan on AWS, the AWS Terraform stack runs. When deployed on GCP, Azure, OCI, or Nebius, the matching stack runs instead. The dependent Helm chart receives the correct database endpoint regardless of cloud provider. ## Next Steps - **[Input Parameters and Output Mapping](https://docs.omnistrate.com/build-guides/terraform-params-outputs/index.md)**: Pass dynamic input variables and map Terraform outputs to other resources - **[Helm and Terraform](https://docs.omnistrate.com/build-guides/helm-charts-terraform/index.md)**: Detailed walkthrough of combining Helm charts with Terraform - **[Custom Terraform Permissions](https://docs.omnistrate.com/getting-started/build-from-terraform/#additional-permissions-for-terraform)**: Configure IAM policies for BYOC deployments # Terraform Overview ## What Is Terraform? Terraform is an infrastructure-as-code (IaC) tool that lets you define and provision cloud infrastructure using declarative configuration files. It supports all major cloud providers — AWS, GCP, Azure, OCI, and Nebius — through a consistent workflow of `plan`, `apply`, and `destroy`. Omnistrate also supports [OpenTofu](https://opentofu.org/) as a drop-in replacement for Terraform. Throughout this guide, references to "Terraform" apply equally to OpenTofu unless stated otherwise. ### Key Concepts **Providers** are plugins that interact with cloud APIs. Each provider (AWS, GCP, Azure, OCI, Nebius) exposes a set of resources you can manage — compute instances, databases, networking, storage, IAM roles, and more. **Resources** are the individual infrastructure components you declare in your Terraform configuration — an S3 bucket, an RDS instance, a VPC, a GCP Pub/Sub topic, etc. **Variables** are input parameters that make your Terraform configuration reusable. You declare variables with types and defaults, then pass values at deployment time. **Outputs** are values exported from your Terraform configuration after `apply` completes — endpoints, ARNs, connection strings, IP addresses. Other components in your Plan can consume these outputs. **State** is the mapping between your declared configuration and real-world infrastructure. Omnistrate manages Terraform state automatically for every deployment. ## Why Use Terraform with Omnistrate? ### Provision Managed Cloud Services Use Terraform to provision managed services — RDS databases, ElastiCache clusters, S3 buckets, DynamoDB tables, GCP Cloud SQL instances, Azure Cosmos DB — as part of your SaaS Product topology. These resources are provisioned alongside your Helm chart, Operator, or container-based deployments. ### Build Multi-Cloud SaaS Products Define cloud-specific Terraform stacks for AWS, GCP, Azure, OCI, and Nebius within a single Plan specification. Omnistrate selects the correct stack based on where the deployment runs, so you can offer your SaaS Product across those clouds from one specification. ### Connect Infrastructure to Application Layers Terraform outputs (endpoints, ARNs, connection strings) can be injected into Helm chart values, environment variables, or other resource configurations. This lets you wire managed cloud services directly into your application without manual configuration. ### Automate Customer-Specific Infrastructure Each customer deployment gets its own isolated Terraform state and resources. Omnistrate handles per-tenant provisioning, updates, and teardown automatically — you define the infrastructure once, and every customer gets their own instance. ## How Terraform Works on Omnistrate ### Terraform as a Resource In Omnistrate, a Terraform stack is defined as a **Resource** within your Plan specification — just like a Helm chart or an Operator. You reference a Git repository containing your Terraform configuration, and Omnistrate executes the stack during the deployment lifecycle. A typical Terraform resource is marked as `internal: true`, meaning it is not directly exposed to your customers. Instead, other resources (Helm charts, Operators) depend on it and consume its outputs. ### Lifecycle Management Omnistrate manages the full Terraform lifecycle: - **Create**: Runs `terraform apply` when a customer deployment is provisioned - **Update**: Runs `terraform apply` with updated variables when configuration changes - **Delete**: Runs `terraform destroy` when a customer deployment is torn down State is managed per deployment, ensuring complete isolation between customers. ### State Backend Management Warning Omnistrate manages the Terraform state backend automatically using Kubernetes secrets. Any custom `backend` configuration you define in your Terraform files (such as S3, GCS, Azure Blob Storage, or any other remote backend) will be removed during deployment. Do not rely on a custom backend block in your Terraform configuration — Omnistrate replaces it with its own managed backend to ensure state consistency and isolation across deployments. ### Control-Plane-Targeted Terraform Not every Terraform stack belongs in the data plane deployment cell. Some stacks manage provider-side resources that should execute in the Omnistrate control plane account instead. For those resources, set `deploymentTarget.account: ControlPlane` on the Terraform resource and define the provider entry that should execute the stack. For example, a control-plane-targeted AWS stack only needs an `aws` entry under `configurationPerCloudProvider` even if the rest of your product deploys to AWS, GCP, Azure, OCI, and Nebius. There is no shared shorthand for per-provider Terraform configuration. Each provider entry is explicit and evaluated independently. ### Provider Authentication Terraform authentication is provider-specific: - For non-Nebius providers, Omnistrate manages execution identities. On AWS, GCP, Azure, and OCI, Omnistrate auto-creates the identity (for example, `omnistrate-terraform-execution-role` on AWS and `omnistrate-tf-` on GCP). Assign the required permissions to the identity. - Nebius uses explicit `serviceAccountID`, `publicKeyID`, and `privateKeyPEM` fields on the `nebius` provider entry. - Prefer `$secret.` for Nebius `privateKeyPEM` so the PEM stays out of source control. ### System Parameter Injection Omnistrate injects system parameters directly into your Terraform templates at deployment time. These provide runtime context such as: - Region and availability zone information - VPC and subnet IDs from the deployment cell - Unique deployment identifiers for resource naming - Cloud provider network configuration For the full list, see [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ## Common Use Cases ### Managed Databases as Dependencies Provision RDS, Cloud SQL, or Azure Database instances through Terraform and pass connection endpoints to your application via output mapping. ### Storage and Messaging Infrastructure Create S3 buckets, GCS buckets, SQS queues, Pub/Sub topics, or Azure Service Bus resources that your application needs at runtime. ### Networking and Security Set up security groups, firewall rules, VPC peering, or private endpoints as part of the deployment — ensuring each customer's infrastructure meets your security requirements. ### IAM and Access Control Create cloud-specific IAM roles, service accounts, or managed identities that your application workloads assume for secure access to cloud services. ## Getting Started The process for building a Terraform-based resource on Omnistrate follows these steps: 1. **Write your Terraform configuration**: Define the cloud resources you need in standard `.tf` files 1. **Store in a Git repository**: Push your Terraform stack to a GitHub repository (public or private) 1. **Define the resource in your Plan specification**: Reference the Git repository and configure per-cloud-provider stacks 1. **Map outputs to dependent resources**: Use output references to pass Terraform outputs to Helm charts or other resources 1. **Build and deploy**: Use the Omnistrate CLI to build your Plan and test in a development environment Here is a minimal example that provisions an S3 bucket and passes its ARN to a Helm chart: ``` name: My SaaS Product deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: infraTerraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/your-terraform-repo.git - name: MyApp dependsOn: - infraTerraform helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-repo chartRepoURL: https://charts.example.com chartValues: s3BucketARN: "{{ $infraTerraform.out.s3_bucket_arn }}" ``` ## Next Steps - **[Multi-Cloud Configuration](https://docs.omnistrate.com/build-guides/terraform-multi-cloud/index.md)**: Configure Terraform stacks for AWS, GCP, Azure, OCI, and Nebius in a single Plan - **[Input Parameters and Output Mapping](https://docs.omnistrate.com/build-guides/terraform-params-outputs/index.md)**: Pass input variables to Terraform and map outputs to other resources - **[Helm and Terraform](https://docs.omnistrate.com/build-guides/helm-charts-terraform/index.md)**: End-to-end example combining Helm charts with Terraform infrastructure - **[System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md)**: Full reference for system parameters available in Terraform templates - **[Getting Started with Terraform](https://docs.omnistrate.com/getting-started/build-from-terraform/index.md)**: Step-by-step tutorial for your first Terraform-based deployment # Input Parameters and Output Mapping ## Overview Terraform resources in Omnistrate support two key integration mechanisms: - **Input parameters**: Pass dynamic values into your Terraform configuration through variables — including system parameters, API parameters, and inline variable overrides - **Output mapping**: Export values from your Terraform state (endpoints, ARNs, IDs) and inject them into dependent resources like Helm charts, Operators, or other Terraform stacks Together, these let you build fully wired, multi-resource SaaS Products where infrastructure provisioned by Terraform is seamlessly connected to application layers. ## Input Parameters ### System Parameters in Terraform Templates Omnistrate injects system parameters directly into your `.tf` files at deployment time. Wrap system parameter references in `{{ }}` within your Terraform configuration: ``` provider "aws" { region = "{{ $sys.deploymentCell.region }}" } resource "aws_security_group" "app_sg" { name = "app-sg-{{ $sys.id }}" vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" ingress { from_port = 5432 to_port = 5432 protocol = "tcp" cidr_blocks = ["{{ $sys.deploymentCell.cidrRange }}"] } } resource "aws_db_subnet_group" "main" { name = "db-subnet-{{ $sys.id }}" subnet_ids = [ "{{ $sys.deploymentCell.privateSubnetIDs[0].id }}", "{{ $sys.deploymentCell.privateSubnetIDs[1].id }}" ] } ``` Commonly used system parameters in Terraform templates: | Parameter | Description | | -------------------------------------------- | --------------------------------------------------------------------------- | | `$sys.id` | Unique deployment identifier — use for naming resources to avoid collisions | | `$sys.deploymentCell.region` | Target deployment region | | `$sys.deploymentCell.cloudProviderNetworkID` | VPC / VNet / Network ID | | `$sys.deploymentCell.cidrRange` | CIDR range of the deployment cell | | `$sys.deploymentCell.publicSubnetIDs[i].id` | Public subnet IDs | | `$sys.deploymentCell.privateSubnetIDs[i].id` | Private subnet IDs | For the complete list, see [System Parameters](https://docs.omnistrate.com/build-guides/system-parameters/index.md). ### Variable Overrides with Inline Values Use the `variablesValuesFileOverride` property to pass Terraform variable values directly from your Plan specification — no separate `.tfvars` file needed in your repository. This is useful for injecting system parameters into Terraform variables. Define variables in your Terraform stack: ``` # variables.tf variable "vpc_id" { type = string description = "VPC ID for the deployment" } variable "region" { type = string description = "AWS region" } variable "instance_type" { type = string default = "db.t3.medium" description = "RDS instance class" } variable "subnet_ids" { type = list(string) description = "Subnet IDs for the database subnet group" } ``` Then override them in your Plan specification: ``` services: - name: dbInfra internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws variablesValuesFileOverride: | vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" instance_type = "db.t3.medium" subnet_ids = [ "{{ $sys.deploymentCell.privateSubnetIDs[0].id }}", "{{ $sys.deploymentCell.privateSubnetIDs[1].id }}" ] gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Tip The `variablesValuesFileOverride` is useful when you want to customize Terraform behavior on the Omnistrate platform without modifying the underlying Terraform code. Your Terraform stack stays generic and reusable while environment-specific values are injected from the Plan specification. ### API Parameters as Terraform Inputs You can connect API parameters — values your customers provide at deployment time — to Terraform variables through the same mechanism. Define an API parameter on the parent resource and reference it in the Terraform variable override: ``` services: - name: MyApp dependsOn: - dbInfra apiParameters: - key: dbInstanceClass description: Database instance size name: Database Instance Class type: String modifiable: true required: false export: true defaultValue: "db.t3.medium" options: - "db.t3.micro" - "db.t3.medium" - "db.t3.large" parameterDependencyMap: dbInfra: dbInstanceClass - name: dbInfra internal: true apiParameters: - key: dbInstanceClass description: Database instance class name: Database Instance Class type: String modifiable: true required: false export: false defaultValue: "db.t3.medium" terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws variablesValuesFileOverride: | instance_type = "{{ $var.dbInstanceClass }}" vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` In this pattern: 1. The customer selects a `dbInstanceClass` when creating a deployment 1. The value is passed through the `parameterDependencyMap` to the `dbInfra` resource 1. The `variablesValuesFileOverride` injects the value into the Terraform variable ### Nebius Terraform Authentication Inputs Nebius Terraform resources use explicit authentication fields on the `nebius` plan entry: ``` services: - name: nebiusInfra internal: true terraformConfigurations: configurationPerCloudProvider: nebius: terraformPath: /terraform/nebius serviceAccountID: serviceaccount-e00vqdp9fskhmmaan8 publicKeyID: publickey-e00h9scsyy9mbefrjf privateKeyPEM: $secret.nebiusTerraformPrivateKey gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git ``` Notes: - `serviceAccountID`, `publicKeyID`, and `privateKeyPEM` are all required for `configurationPerCloudProvider.nebius`. - Prefer `$secret.` for `privateKeyPEM`; Omnistrate resolves the secret before Terraform runs. - These fields apply to the Terraform resource itself. They are not written through `variablesValuesFileOverride`. - When Nebius auth is specified, Omnistrate uses it for that Terraform resource instead of the default host-cluster Terraform identity. ## Output Mapping ### Defining Terraform Outputs In your Terraform stack, define outputs for any values you want to expose to other resources: ``` # outputs.tf output "database_endpoint" { value = aws_db_instance.main.endpoint description = "RDS instance endpoint" } output "database_port" { value = aws_db_instance.main.port description = "RDS instance port" } output "s3_bucket_arn" { value = aws_s3_bucket.data.arn description = "S3 bucket ARN for data storage" } output "connection_details" { value = { endpoint = aws_db_instance.main.endpoint port = aws_db_instance.main.port database = aws_db_instance.main.db_name } description = "Structured connection details" } output "cache_endpoint" { value = aws_elasticache_cluster.cache.cache_nodes[0].address sensitive = true } ``` Omnistrate automatically captures all values from the Terraform `output` block after each `apply`. ### Exporting Selected Terraform Outputs All Terraform outputs are available to dependent resources through `{{ $.out. }}`. If you also want a selected output to surface as an exported field on the Terraform resource itself, declare it under `requiredOutputs`. ``` services: - name: infraTerraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git requiredOutputs: - key: database_endpoint exported: true - key: cache_endpoint exported: true ``` Use this when you want the output to be visible as part of the Terraform resource details in Omnistrate, not just consumed by another resource in the DAG. Note `requiredOutputs` does not replace normal output mapping. Dependent resources still consume Terraform outputs using `{{ $.out. }}`. ### Referencing Outputs in Dependent Resources Use the `{{ $.out. }}` syntax to inject Terraform outputs into other resources. The `` is the `name` of the Terraform resource as defined in your Plan specification. #### In Helm Chart Values ``` services: - name: infraTerraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git - name: MyApp dependsOn: - infraTerraform helmChartConfiguration: chartName: my-app chartVersion: 1.0.0 chartRepoName: my-charts chartRepoURL: https://charts.example.com chartValues: database: host: "{{ $infraTerraform.out.database_endpoint }}" port: "{{ $infraTerraform.out.database_port }}" storage: bucketArn: "{{ $infraTerraform.out.s3_bucket_arn }}" cache: endpoint: "{{ $infraTerraform.out.cache_endpoint }}" ``` #### Accessing Nested Output Values When a Terraform output returns a structured object, you can access nested properties with dot notation: ``` # Terraform output output "connection_details" { value = { endpoint = aws_db_instance.main.endpoint port = aws_db_instance.main.port } } ``` ``` # Plan specification reference chartValues: dbHost: "{{ $infraTerraform.out.connection_details.endpoint }}" dbPort: "{{ $infraTerraform.out.connection_details.port }}" ``` #### In Operator CRD Configuration Note This example uses the legacy `operatorCRDConfiguration.template` field to show Terraform outputs flowing into an operator resource. For new operator-backed services, prefer `systemWorkflows` for lifecycle resources, readiness, failure conditions, and outputs. See [Build with Kubernetes Operators](https://docs.omnistrate.com/build-guides/operators/index.md). ``` services: - name: infraTerraform internal: true terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws gitConfiguration: reference: refs/heads/main repositoryUrl: https://github.com/your-org/infra-repo.git - name: DatabaseOperator dependsOn: - infraTerraform operatorCRDConfiguration: template: | apiVersion: db.example.com/v1 kind: Database spec: externalEndpoint: "{{ $infraTerraform.out.database_endpoint }}" bucketArn: "{{ $infraTerraform.out.s3_bucket_arn }}" ``` ### Outputs Across Multi-Cloud Stacks When you define Terraform stacks for multiple cloud providers, keep output names consistent so dependent resources work without cloud-specific logic: **AWS outputs:** ``` output "database_endpoint" { value = aws_db_instance.main.endpoint } ``` **GCP outputs:** ``` output "database_endpoint" { value = google_sql_database_instance.main.private_ip_address } ``` **Azure outputs:** ``` output "database_endpoint" { value = azurerm_postgresql_flexible_server.main.fqdn } ``` The dependent resource references `{{ $infraTerraform.out.database_endpoint }}` regardless of which cloud provider is in use. ## End-to-End Example This complete example demonstrates a multi-resource Plan with a Terraform stack provisioning cloud infrastructure, API parameters exposed to customers, and output mapping to a Helm chart application: ``` name: Full Stack SaaS deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role services: - name: cloudInfra internal: true apiParameters: - key: dbInstanceClass description: Database instance class name: Database Instance Class type: String modifiable: true required: false export: false defaultValue: "db.t3.medium" - key: storageSize description: Database storage size in GB name: Storage Size (GB) type: String modifiable: true required: false export: false defaultValue: "20" terraformConfigurations: configurationPerCloudProvider: aws: terraformPath: /terraform/aws variablesValuesFileOverride: | vpc_id = "{{ $sys.deploymentCell.cloudProviderNetworkID }}" region = "{{ $sys.deploymentCell.region }}" instance_class = "{{ $var.dbInstanceClass }}" storage_size = {{ $var.storageSize }} subnet_ids = [ "{{ $sys.deploymentCell.privateSubnetIDs[0].id }}", "{{ $sys.deploymentCell.privateSubnetIDs[1].id }}" ] resource_prefix = "saas-{{ $sys.id }}" gitConfiguration: reference: refs/tags/v1.0.0 repositoryUrl: https://github.com/your-org/infra-repo.git - name: WebApp dependsOn: - cloudInfra network: ports: - 8080 compute: instanceTypes: - apiParam: instanceType cloudProvider: aws helmChartConfiguration: chartName: web-app chartVersion: 2.0.0 chartRepoName: my-charts chartRepoURL: https://charts.example.com chartValues: config: databaseUrl: "{{ $cloudInfra.out.database_endpoint }}" databasePort: "{{ $cloudInfra.out.database_port }}" s3BucketArn: "{{ $cloudInfra.out.s3_bucket_arn }}" apiParameters: - key: instanceType description: Compute instance type name: Instance Type type: String modifiable: true required: false export: true defaultValue: "t4g.small" - key: dbInstanceClass description: Database instance size name: Database Instance Class type: String modifiable: true required: false export: true defaultValue: "db.t3.medium" options: - "db.t3.micro" - "db.t3.medium" - "db.t3.large" parameterDependencyMap: cloudInfra: dbInstanceClass - key: storageSize description: Database storage in GB name: Storage Size (GB) type: String modifiable: true required: false export: true defaultValue: "20" parameterDependencyMap: cloudInfra: storageSize ``` In this example: 1. **`cloudInfra`** is an internal Terraform resource that provisions an RDS database and S3 bucket 1. **`WebApp`** is the customer-facing Helm chart that depends on `cloudInfra` 1. API parameters (`dbInstanceClass`, `storageSize`) are defined on `WebApp` and mapped to `cloudInfra` through `parameterDependencyMap` 1. Terraform outputs (`database_endpoint`, `database_port`, `s3_bucket_arn`) are injected into the Helm chart values 1. System parameters provide the VPC, region, and subnet context to Terraform ## Direct vs Transitive Parameter Mapping `parameterDependencyMap` is direct between a parent resource and its immediate child. It is not transitive. If `App` depends on both `Infra` and `SecureInfra`, and `SecureInfra` also depends on `Infra`, then: - Parameters that `Infra` needs must be mapped directly from `App` to `Infra` - Parameters mapped from `App` to `SecureInfra` are not automatically forwarded to `Infra` - Outputs from `Infra` can still be consumed by `SecureInfra` through `{{ $Infra.out. }}` Example: ``` services: - name: Infra internal: true apiParameters: - key: access_type name: Access Type type: String export: false required: false modifiable: true defaultValue: public - name: SecureInfra internal: true dependsOn: - Infra apiParameters: - key: tenant_name name: Tenant Name type: String export: false required: false modifiable: true defaultValue: tenant-a - name: App dependsOn: - Infra - SecureInfra apiParameters: - key: accessType name: Access Type type: String export: true required: false modifiable: true defaultValue: public parameterDependencyMap: Infra: access_type - key: tenantName name: Tenant Name type: String export: true required: false modifiable: true defaultValue: tenant-a parameterDependencyMap: SecureInfra: tenant_name ``` ## Best Practices ### Consistent Output Names Use the same output key names across your AWS, GCP, Azure, OCI, and Nebius Terraform stacks. This ensures dependent resources work without cloud-specific branching. ### Use Git Tags for Versioning Reference specific Git tags (`refs/tags/v1.0.0`) instead of branches for production Plans. This ensures deployments are reproducible and upgrades are intentional. ### Unique Resource Names Always include `$sys.id` in cloud resource names to prevent naming collisions between customer deployments: ``` resource "aws_db_instance" "main" { identifier = "db-{{ $sys.id }}" ... } ``` ### Mark Terraform Resources as Internal Terraform resources are infrastructure dependencies — mark them as `internal: true` so they are not directly exposed in your customer-facing API. Customers interact with the parent resource that depends on the Terraform stack. ### Sensitive Outputs Mark Terraform outputs as `sensitive = true` for values like passwords or connection strings that should not appear in logs: ``` output "db_password" { value = random_password.db.result sensitive = true } ``` ## Next Steps - **[Terraform Overview](https://docs.omnistrate.com/build-guides/terraform-overview/index.md)**: Understand how Terraform integrates with the Omnistrate platform - **[Multi-Cloud Configuration](https://docs.omnistrate.com/build-guides/terraform-multi-cloud/index.md)**: Configure per-cloud-provider Terraform stacks - **[Helm and Terraform](https://docs.omnistrate.com/build-guides/helm-charts-terraform/index.md)**: Complete walkthrough combining Helm with Terraform - **[API Parameters](https://docs.omnistrate.com/build-guides/api-params/index.md)**: Full guide on defining customer-facing parameters - **[Resource Dependencies](https://docs.omnistrate.com/build-guides/dependencies/index.md)**: How resource ordering and parameter mapping works # Terraform Troubleshooting For the shared debugging workflow, start with [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md). It explains Debug Events, instance debug, and the relationship between workflow state and live resource state. ## Terraform Execution Model and Debugging Expectations Omnistrate executes Terraform and OpenTofu autonomously as part of the instance workflow. The platform does not support a manual review or approval gate for `terraform plan` before `terraform apply`. The recommended operating model is: 1. Test the change in a lower environment first. 1. Pin the Terraform source with a Git tag or commit SHA. 1. Promote the same tested artifacts to production. ## Troubleshooting Terraform Deployments If a Terraform-based deployment fails, start from [Debug Events](https://docs.omnistrate.com/operate-guides/troubleshooting/#debug-events) to identify whether the failure happened during artifact rendering, `terraform plan`, `terraform apply`, or a later dependent resource. Then inspect the Terraform resource’s operation history with: ``` omnistrate-ctl instance debug ``` Use the instance debug output to investigate: 1. Rendered Terraform files after Omnistrate substitutes system parameters, API parameters, secrets, and dependency outputs. 1. Captured `terraform plan` output to confirm Terraform intended to create, update, or delete the resources you expected. 1. Terraform apply logs for cloud provider errors such as missing IAM permissions, quota limits, unavailable regions or SKUs, invalid networking inputs, and resource name conflicts. 1. Provider configuration and credentials, especially when using custom Terraform permissions, BYOC onboarding, control-plane-targeted Terraform, or Nebius service-account authentication. 1. Terraform outputs that feed downstream Helm, Operator, Kustomize, or Terraform resources. The shared [Terraform Troubleshooting Checklist](https://docs.omnistrate.com/operate-guides/troubleshooting/#terraform-troubleshooting-checklist) contains the same checks alongside the other deployment strategies. If the rendered artifacts are wrong, update the Plan spec, Terraform source, parameter mappings, or Git reference and publish a new Plan version before retrying. If the plan looks correct but apply fails, fix the cloud-side permission, quota, naming, or provider issue and retry the workflow when the same released artifacts are still valid. ## Restarting a Workflow vs Publishing a New Version Restarting a failed workflow is best for transient issues when you want to retry the same released artifacts again. If you changed the Plan specification, a `parameterDependencyMap`, the referenced Git branch/tag/commit, or Terraform or Helm artifacts that must be re-rendered, publish a new Plan version and trigger a fresh workflow instead of restarting the old one. As a rule of thumb, assume a workflow restart retries the artifacts already captured for that workflow. For deterministic behavior, pin the Git source to a tag or commit SHA. # Runtime Guides # Autoscaling ## Overview Omnistrate enables your Resources to have advanced autoscaling capabilities using custom metrics. You can configure autoscaling to dynamically adjust the number of replicas based on load, helping you optimize infrastructure costs while maintaining performance. Omnistrate offers advanced autoscaling that works for your stateful systems. Here are some specific features in addition to simply autoscaling based on the load: - Scale machines on demand based on custom metrics, not only the load metric average. A common problem occurs if one node is misbehaving and has high load; the solution in those cases is not to always add more resources - Run specific operations as part of the scale operation. Omnistrate allows you to run custom code as part of the scale operation, for example, to call rebalance, update cluster metadata, or rotate certificates - Automatically coordinate different operations on your Resources to prevent complex operations from interfering with each other Note To scale down to zero, use [Custom Autoscaling](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md) or [Serverless](https://docs.omnistrate.com/runtime-guides/serverless/index.md) capabilities. ## Configuring auto-scaling You can enable autoscaling for one or more Resources by adding the autoscaling configuration under `x-omnistrate-capabilities`. Here is an example using compose spec: ``` services: app: x-omnistrate-capabilities: autoscaling: maxReplicas: 5 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 ``` ## Configuration parameters - **maxReplicas**: Maximum number of replicas to scale up to - **minReplicas**: Minimum number of replicas to maintain. The value must be between [1, maxReplicas]. To scale down to zero, use [Custom Autoscaling](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md) or [Serverless](https://docs.omnistrate.com/runtime-guides/serverless/index.md) - **idleMinutesBeforeScalingDown**: Duration in minutes to wait before scaling down when idle - **idleThreshold**: Metric threshold value below which the system is considered idle and will scale down after `idleMinutesBeforeScalingDown` - **overUtilizedMinutesBeforeScalingUp**: Duration in minutes to wait before scaling up when over-utilized - **overUtilizedThreshold**: Metric threshold value above which the system is considered over-utilized and will scale up after `overUtilizedMinutesBeforeScalingUp` - **scalingMetric**: Optional custom metric configuration. By default, Omnistrate uses CPU load metrics, but you can configure your own custom metrics Note For custom metrics, Omnistrate leverages Prometheus endpoints and requires your application to emit the required metrics. Metric source limitations Native autoscaling expects infrastructure metrics or a Prometheus endpoint exposed by your application. It does not scrape arbitrary authenticated HTTP APIs directly. If your scaling signal comes from a queue API, requires authentication, or needs custom business logic, use [Custom Autoscaling](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md) instead. ## Configuring auto-scaling with custom metrics You can enable autoscaling for one or more Resources by adding the autoscaling configuration under `x-omnistrate-capabilities` and defining metrics configuration. Here is an example using compose spec: ``` services: app: x-omnistrate-capabilities: autoscaling: maxReplicas: 5 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 scalingMetric: metricEndpoint: "http://localhost:9187/metrics" metricName: "custom_metric_name" metricLabelName: "application_name" metricLabelValue: "custom_metric_label" ``` Use this mode when your application can expose a Prometheus metric locally. If your signal comes from an external system such as a queue API or a protected endpoint, either export that signal as a Prometheus metric first or switch to [Custom Autoscaling](https://docs.omnistrate.com/runtime-guides/custom-autoscaling/index.md). ## Service Quotas Your SaaS Product leverages some basic cloud resources to host your customer's workloads. When you onboard your cloud account, please make sure to check the following quotas in your cloud account and request an increase if needed: ### AWS 1. **Number of nodegroups per EKS cluster** - Recommended: 100 - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/eks/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **Running On-Demand Standard (A, C, D, H, I, M, R, T, Z) instances** - Recommended: 16 minimum + adjust for your workload - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/ec2/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of customer networks you plan to host in your account and increase this limit accordingly 1. **NAT gateways per Availability Zone** - Recommended: Workload dependent (if you plan to isolate your customers in their own VPCs) - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/vpc/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **EC2-VPC Elastic IPs** - Recommended: Workload dependent (if you plan to isolate your customers in their own VPCs) - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/ec2/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **Running On-Demand G and VT instances (for GPU workloads)** - Recommended: Workload dependent - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/ec2/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of customer networks you plan to host in your account and increase this limit accordingly 1. **Running On-Demand P instances (for GPU workloads)** - Recommended: Workload dependent - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/ec2/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of customer networks you plan to host in your account and increase this limit accordingly 1. **Nodes per managed node group** - Recommended: Workload dependent - [Request an increase](https://us-east-2.console.aws.amazon.com/servicequotas/home/services/eks/quotas) - Make sure to do this for each region you plan to operate your SaaS in ### GCP 1. **In-use regional external IPv4 addresses** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **CPUs** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **Persistent Disk SSD (GB)** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in 1. **Networks** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of customer networks you plan to host in your account and increase this limit accordingly 1. **Routers** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of customer networks you plan to host in your account and increase this limit accordingly 1. **Firewall rules** - Recommended: Workload dependent - [Request an increase](https://console.cloud.google.com/apis/api/compute.googleapis.com/quotas) - Make sure to do this for each region you plan to operate your SaaS in - Also consider the number of Plans and the complexity (number of Resources) of each Plan you plan to offer and increase this limit accordingly ### Azure 1. **Total Regional vCPUs** - Recommended: Workload dependent - [Request an increase](https://portal.azure.com/#view/Microsoft_Azure_Capacity/QuotaMenuBlade/~/myQuotas) - Make sure to do this for each region you plan to operate your SaaS in - Default: 65 1. **Standard Basv2 Family vCPUs** - Recommended: Workload dependent - [Request an increase](https://portal.azure.com/#view/Microsoft_Azure_Capacity/QuotaMenuBlade/~/myQuotas) - Make sure to do this for each region you plan to operate your SaaS in - This is the default instance family used by Omnistrate. If you are planning to use a different instance family, please adjust its limits instead. - Default: 65 1. **Public IP Addresses** - Recommended: Workload dependent - [Request an increase](https://portal.azure.com/#view/Microsoft_Azure_Capacity/QuotaMenuBlade/~/myQuotas) - Make sure to do this for each region you plan to operate your SaaS in - Each resource in the public subnet using dedicated tenancy will consume a public IP address. - Default: 100 ### OCI 1. **Compute Core Count** - Recommended: Workload dependent - [Request an increase](https://cloud.oracle.com/limits) via the Service Limits page in OCI Console - Make sure to do this for each region you plan to operate your SaaS in 1. **Virtual Cloud Network (VCN) Count** - Recommended: Workload dependent - [Request an increase](https://cloud.oracle.com/limits) via the Service Limits page in OCI Console - Make sure to do this for each region you plan to operate your SaaS in 1. **Load Balancer Count** - Recommended: Workload dependent - [Request an increase](https://cloud.oracle.com/limits) via the Service Limits page in OCI Console - Make sure to do this for each region you plan to operate your SaaS in 1. **Block Volume Size (GB)** - Recommended: Workload dependent - [Request an increase](https://cloud.oracle.com/limits) via the Service Limits page in OCI Console - Make sure to do this for each region you plan to operate your SaaS in # Custom Autoscaling ## Overview Custom autoscaling enables you to implement programmatic control over your Resource scaling decisions when standard metric-based autoscaling doesn't meet your requirements. Unlike Omnistrate's built-in autoscaling that relies on predefined infrastructure or application metrics, custom autoscaling allows you to define completely custom scaling logic based on any business rules, external systems, or complex decision-making processes. When you enable custom autoscaling for a Resource, Omnistrate provides a local sidecar API that allows your controller to query capacity information and trigger scaling operations. This gives you the flexibility to implement scaling logic in any programming language while leveraging Omnistrate's capacity management capabilities. ## When to use custom autoscaling Use Omnistrate's native [metric-based autoscaling](https://docs.omnistrate.com/runtime-guides/autoscaling/index.md) when: - Scaling based on CPU, memory, or application metrics (queue length, request rate, etc.) meets your needs - You want a fully managed, zero-code autoscaling solution - Simple threshold-based scaling rules are sufficient for your workload - Your application can expose the required signal as a Prometheus metric - You don't need custom business logic or complex decision-making Use custom autoscaling when: - You need to implement complex business logic that combines multiple conditions or rules - You want predictive scaling based on ML models, time series forecasting, or historical patterns - You require custom scheduling with complex rules (holidays, events, multi-stage rollouts) - You need to integrate external systems into your scaling decisions - You want to implement custom cooldown strategies or multi-resource coordination - You need programmatic control over scaling with language-specific libraries or frameworks - Standard metric-based thresholds don't capture your scaling requirements - You need to scale down to zero replicas (custom autoscaling supports scaling to zero) - Your scaling signal comes from an authenticated API, queue service, or external system that cannot be scraped directly as a Prometheus endpoint ## Configuring custom autoscaling To enable custom autoscaling for a Resource, set `policyType: custom` in the autoscaling configuration under `x-omnistrate-capabilities`. Here is an example using compose spec: ``` services: worker: x-omnistrate-capabilities: autoscaling: policyType: custom maxReplicas: 6 minReplicas: 1 ``` ### Configuration parameters - **policyType**: Set to `custom` to enable custom autoscaling - **maxReplicas**: Maximum number of replicas to scale up to - **minReplicas**: Minimum number of replicas to maintain. Can be set to 0 to allow scaling down to zero ## Sidecar API When you deploy a SaaS Product with custom autoscaling enabled, Omnistrate automatically provides a local sidecar API. This API enables your custom controller to query capacity information and trigger scaling operations. This is the recommended path when native autoscaling is too restrictive because your metrics source is not a simple Prometheus endpoint or when the scale decision must include custom business logic. The local sidecar API is available at: ``` http://127.0.0.1:49750 ``` Note The sidecar API is only available when running on Omnistrate. It is not available for local development outside of the Omnistrate platform. ### API endpoints Single-resource API endpoints are accessible at: ``` http://127.0.0.1:49750/resource/{resourceAlias} ``` Where `{resourceAlias}` is the Resource key from your compose specification file. Multi-resource API endpoints are accessible at: ``` http://127.0.0.1:49750/resources/capacity/{operation} ``` Where `{operation}` is either `add` or `remove`. #### Get current capacity Retrieves the current capacity and status information for a Resource. **Endpoint**: `GET /resource/{resourceAlias}/capacity` **Response**: ``` { "instanceId": "string", "resourceId": "string", "resourceAlias": "string", "status": "ACTIVE|STARTING|PAUSED|FAILED|UNKNOWN", "currentCapacity": 5, "lastObservedTimestamp": "2025-11-05T12:34:56.789Z" } ``` **Status values**: - `ACTIVE` - Resource is running and ready - `STARTING` - Resource is starting up - `PAUSED` - Resource is paused - `FAILED` - Resource has failed - `UNKNOWN` - Status cannot be determined #### Add capacity Adds capacity units to a Resource. **Endpoint**: `POST /resource/{resourceAlias}/capacity/add` **Request body**: ``` { "capacityToBeAdded": 2 } ``` **Response**: ``` { "instanceId": "string", "resourceId": "string", "resourceAlias": "string" } ``` #### Remove capacity Removes capacity units from a Resource. **Endpoint**: `POST /resource/{resourceAlias}/capacity/remove` **Request body**: ``` { "capacityToBeRemoved": 1 } ``` **Response**: ``` { "instanceId": "string", "resourceId": "string", "resourceAlias": "string" } ``` #### Add capacity to multiple Resources Adds capacity units to multiple Resources in the same instance. **Endpoint**: `POST /resources/capacity/add` **Request body**: ``` { "resources": [ { "resourceAlias": "api", "capacityToBeAdded": 2 }, { "resourceAlias": "worker", "capacityToBeAdded": 1 } ] } ``` **Response**: ``` { "instanceId": "string", "resources": [ { "instanceId": "string", "resourceId": "string", "resourceAlias": "api" }, { "instanceId": "string", "resourceId": "string", "resourceAlias": "worker" } ] } ``` #### Remove capacity from multiple Resources Removes capacity units from multiple Resources in the same instance. **Endpoint**: `POST /resources/capacity/remove` **Request body**: ``` { "resources": [ { "resourceAlias": "api", "capacityToBeRemoved": 1 }, { "resourceAlias": "worker", "capacityToBeRemoved": 2 } ] } ``` **Response**: ``` { "instanceId": "string", "resources": [ { "instanceId": "string", "resourceId": "string", "resourceAlias": "api" }, { "instanceId": "string", "resourceId": "string", "resourceAlias": "worker" } ] } ``` Note The existing single-resource endpoints remain supported for backward compatibility. Use the multi-resource endpoints when a scaling decision needs to add or remove capacity across multiple Resources in the same instance. Grouped capacity requests are applied as one instance-scoped scaling operation. Omnistrate honors the Resource dependency tree during orchestration, so dependent Resources are updated in dependency order while independent Resources can be updated in parallel. For grouped requests: - Each item in `resources` must include a valid `resourceAlias` from the compose specification file. - Each item must include the capacity field that matches the endpoint: `capacityToBeAdded` for `/resources/capacity/add` or `capacityToBeRemoved` for `/resources/capacity/remove`. - Capacity values must be positive integers. - Duplicate `resourceAlias` values in the same request are rejected. - The request can only target Resources in the same instance. ## Implementation best practices ### Implement cooldown periods Avoid rapid successive scaling operations by implementing a cooldown period between scaling actions. A recommended cooldown period is 5 minutes (300 seconds). ``` Last Scale Action → Wait 5 minutes → Next Scale Action ``` ### Wait for ACTIVE state Always wait for a Resource to reach `ACTIVE` state before performing the next scaling operation: 1. Check current status via GET capacity endpoint 1. If status is `STARTING`, poll until it becomes `ACTIVE` 1. If status is `FAILED`, handle the error appropriately 1. Only proceed with scaling when status is `ACTIVE` ### Implement step-based scaling Scale gradually by adding or removing a fixed number of units per operation: - Start with small steps (e.g., 1-2 units) - Repeat operations until reaching target capacity - Allow cooldown between steps ### Respect scaling limits Always validate that your scaling requests stay within the `minReplicas` and `maxReplicas` limits you defined in the Resource specification. Attempting to scale beyond these limits will fail. Your custom controller should: - Check current capacity before calculating scaling targets - Ensure target capacity is between `minReplicas` and `maxReplicas` - Handle edge cases when already at minimum or maximum capacity ### Use retries with exponential backoff Network requests can fail temporarily. Implement retry logic: - Retry failed requests - Use exponential backoff - Set appropriate timeouts ### Handle errors gracefully - Check HTTP status codes (expect 200 for success) - Parse error responses - Log errors for debugging ## Example implementation For a complete working example of single-Resource custom autoscaling, refer to the [Custom Auto Scaling Example](https://github.com/omnistrate-community/custom-auto-scaling-example) repository. This repository provides: - An example Go implementation with cooldown management and state handling - HTTP API for triggering scaling operations and checking status - Docker containerization for easy deployment - Complete implementation examples in multiple programming languages - Common scaling patterns and best practices For a complete working example that scales multiple Resources from one decision loop, refer to the [Multi-Resource Auto Scaling Example](https://github.com/omnistrate-community/multi-resource-auto-scaling-example) repository. This example uses the multi-resource capacity endpoints and configures the target Resources with `AUTOSCALER_TARGET_RESOURCES`, such as `api,worker`. ## Querying metrics with Prometheus endpoint When you deploy a SaaS Product with Omnistrate, [Prometheus](https://prometheus.io/) endpoint is automatically provided. You can use the endpoint to query system and application metrics. The metrics endpoint is available at: ``` http://127.0.0.1:49751/metrics ``` Note The Prometheus endpoint is only available when running on Omnistrate. It is not available for local development outside of the Omnistrate platform. You can use Prometheus client libraries available for your programming language (such as the official [Prometheus client libraries](https://prometheus.io/docs/instrumenting/clientlibs/) for Go, Java, Python, Ruby, and others) to query the endpoint. ### Example metrics output You can query the metrics endpoint using curl or any HTTP client: ``` curl http://127.0.0.1:49751/metrics ``` The endpoint returns metrics in Prometheus text exposition format: ``` # HELP cpu_usage Current CPU usage # TYPE cpu_usage gauge cpu_usage{customer_visible="true",service_provider_visible="true"} 7.629704984760052 # HELP disk_ops_per_sec Current disk IOPS # TYPE disk_ops_per_sec gauge disk_ops_per_sec{customer_visible="true",disk="/app/storage",service_provider_visible="true",type="read"} 0 disk_ops_per_sec{customer_visible="true",disk="/app/storage",service_provider_visible="true",type="write"} 0 # HELP disk_throughput_bytes_per_sec Disk throughput in bytes per second # TYPE disk_throughput_bytes_per_sec gauge disk_throughput_bytes_per_sec{customer_visible="true",disk="/app/storage",service_provider_visible="true",type="read"} 0 disk_throughput_bytes_per_sec{customer_visible="true",disk="/app/storage",service_provider_visible="true",type="write"} 0 # HELP disk_usage_percent Current disk usage as a percentage of total disk space # TYPE disk_usage_percent gauge disk_usage_percent{customer_visible="true",path="/app/storage",service_provider_visible="true"} 0.006691028941233667 disk_usage_percent{customer_visible="true",path="/var/log/app",service_provider_visible="true"} 0.1497328281402588 # HELP load_avg Load average # TYPE load_avg gauge load_avg{customer_visible="true",period="15min",service_provider_visible="true"} 0 load_avg{customer_visible="true",period="1min",service_provider_visible="true"} 0 load_avg{customer_visible="true",period="5min",service_provider_visible="true"} 0 # HELP mem_total_bytes Total memory in bytes # TYPE mem_total_bytes gauge mem_total_bytes{customer_visible="true",service_provider_visible="true"} 4.294967296e+09 # HELP mem_usage_bytes Current memory usage in bytes # TYPE mem_usage_bytes gauge mem_usage_bytes{customer_visible="true",service_provider_visible="true"} 1.652563968e+09 # HELP mem_usage_percent Current memory usage as a percentage of total memory # TYPE mem_usage_percent gauge mem_usage_percent{customer_visible="true",service_provider_visible="true"} 38.47675323486328 # HELP system_uptime_seconds System uptime in seconds # TYPE system_uptime_seconds gauge system_uptime_seconds{customer_visible="true",service_provider_visible="true"} 6725 ``` ### Available metrics The metrics endpoint provides the following system metrics: - **cpu_usage**: Current CPU usage percentage - **disk_ops_per_sec**: Disk IOPS for read and write operations - **disk_throughput_bytes_per_sec**: Disk throughput in bytes per second - **disk_usage_percent**: Disk usage as a percentage of total disk space - **load_avg**: System load average (1min, 5min, 15min periods) - **mem_total_bytes**: Total available memory in bytes - **mem_usage_bytes**: Current memory usage in bytes - **mem_usage_percent**: Memory usage as a percentage of total memory - **system_uptime_seconds**: System uptime in seconds ### Using metrics for custom autoscaling You can incorporate these metrics into your custom autoscaling logic by: 1. Periodically querying the metrics endpoint from your custom controller 1. Parsing the Prometheus text format to extract relevant metric values 1. Implementing scaling decisions based on metric thresholds or patterns 1. Combining multiple metrics for sophisticated scaling logic For example, you might scale up when CPU usage exceeds 80% and memory usage exceeds 70% simultaneously, or implement predictive scaling based on historical metric patterns. # Custom sidecars Sidecars allow you to enhance your SaaS product's functionality without changing its core service. Think of them as add-on containers that run alongside your main application container. You can: 1. Add new features or capabilities by attaching sidecar containers 1. Use existing open-source implementations as sidecars 1. Deploy your own custom implementations as sidecars 1. Run multiple sidecars with each resource instance The 'SIDECARS' resource capability makes this process simple, letting you expand your service's features in a modular way while maintaining isolation from the main application. For example, you might add a sidecar to handle custom logging, or monitoring, while your main service container focuses on core business logic Most basic configuration requires only an image with a version to be provided. Here is an example using compose spec: ``` x-omnistrate-capabilities: sidecars: tooling: # example sidecar name imageNameWithTag: "busybox:stable" ``` With this configuration, your additional sidecar will inherit environment variables, config maps, and security context from your main resource. Note Pod volumes (including logs) will be automatically mounted to make them accessible for each sidecar container. ## Overriding security context In some cases, sidecar might require a specific user and/or group to execute. If that is your case, you can specify them within the configuration in the spec file as follows: ``` x-omnistrate-capabilities: sidecars: tooling: # example sidecar name imageNameWithTag: "busybox:stable" securityContext: runAsUser: 10 runAsGroup: 99 runAsNonRoot: true capabilities: add: - SYS_RESOURCE ``` ## Providing custom limits Omnistrate will set default resource limits for each sidecar. If you find default limits inappropriate, you can provide custom limits in spec configuration: ``` x-omnistrate-capabilities: sidecars: tooling: # example sidecar name imageNameWithTag: "busybox:stable" resourceLimits: cpu: "250m" memory: "256Mi" ``` For more information about resource limits, see [Kubernetes CPU limits documentation](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#meaning-of-cpu), or [Kubernetes memory limits documentation](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#meaning-of-memory). ## Specifying an entry point Not all containers come with an entry point preconfigured. When configuring the custom sidecar you might need to provide the command and arguments that ensure the sidecar performs its intended task and remains running continuously. Here is an example of a command, whose only purpose is to keep the container running: ``` x-omnistrate-capabilities: sidecars: tooling: # example sidecar name imageNameWithTag: "busybox:stable" command: - "/bin/sh" - "-c" - "while true; do echo hello; sleep 10; done" ``` Extra arguments can be passed by providing values to the "args" property: ``` x-omnistrate-capabilities: sidecars: my-sidecar: # example sidecar name imageNameWithTag: "private/my-image:2.1.3" command: - "/bin/init-command" args: - "setup" - "{{ $sys.id }}" ``` # Custom Tagging Custom (Cost) tagging is one of the features offered by Omnistrate that allows users to organize and manage their cloud resources efficiently. Tags are key-value pairs attached to cloud resources, providing users with a flexible and customizable way to categorize, track, and control their infrastructure components. ### Key Features: - Flexible Tagging: Omnistrate supports the creation and assignment of custom tags to a wide range of cloud resources, including virtual machines, storage buckets, databases, and networking components. - Customizable Tag Structures: Users can define their own tag structures based on their organization's specific needs and workflows. Tags can be organized hierarchically to reflect different attributes such as environment (e.g., production, staging, development), application, department, cost center, or project. - Resource Grouping and Visibility: Tags enable users to group related resources together based on common attributes. This grouping facilitates easier resource management, cost allocation, and access control. Users can quickly identify and visualize resources belonging to specific categories or projects using tag-based filtering and searching capabilities. - Cost Allocation and Budgeting: Infrastructure tagging plays a crucial role in cost allocation and budget tracking by providing granular insights into resource usage and spending. Users can analyze cost breakdowns by tag values and allocate expenses accurately across different teams, projects, or departments. - Policy Enforcement and Governance: Tags serve as valuable metadata for implementing policy enforcement and governance measures within the cloud environment. Users can define and enforce tagging policies to ensure compliance with organizational standards, security requirements, and regulatory guidelines. ### How to use tagging Users can create custom tags directly within the Omnistrate dashboard as one of the capabilities or through [API](https://api.omnistrate.cloud/docs/external/#tag/infra-config-api/operation/infra-config-api#CreateInfraConfig). Tags can be defined with specific key-value pairs to represent different attributes or classifications. Once created, tags can be assigned to individual resources during the provisioning process or retrospectively to existing resources. Users have the flexibility to assign multiple tags to a single resource and modify tag assignments as needed. ### Conclusion: Infrastructure tagging is a fundamental aspect of cloud resource management and optimization. With Omnistrate's comprehensive tagging capabilities, users can effectively organize, monitor, and govern their cloud infrastructure while gaining valuable visibility and control over their environment. By harnessing the power of tags, organizations can streamline operations, improve cost efficiency, and enhance overall governance practices within their cloud ecosystem. # Customer networks Customer networks are an advanced Hosted SaaS feature that allows your Customers to define network partitioning on a dedicated stack while keeping the service deployed through Hosted SaaS. This provides complete isolation through a dedicated stack that is not shared with other Customers. It also enables you to set up a private network path for your customer to connect to the service while keeping the option of provisioning the service as self-served. Note This feature is available for Hosted SaaS only. BYOC always provides isolation by provisioning the services in the customer account directly. ## Enabling Customer Networks Customer Provided Networks can only be disabled (default) or enabled. When enabled, your Customers will have to choose or create a new Custom Network whenever creating a new instance. Instances created with the same Custom Network will be co-located on same Deployment Cell and will share the same network interface. A Customer Network is an abstract concept that defines a CIDR range for a specific Cloud Provider and a Region. Because each such network is owned by certain customer, instances belonging to different Customers within the Plan are never co-located on Deployment Cells with instances from other Customers. Use RFC1918 address space for customer networks, such as `10.0.0.0/8`, `172.16.0.0/12`, or `192.168.0.0/16`. Non-RFC1918 private-network examples are not recommended and can create routing or peering surprises in enterprise environments. Warning Enabling this feature can potentially increase your infrastructure cost significantly as it can result in additional Kubernetes Host Clusters being provisioned, one for each Customer defined Custom Network. Consider this when defining the Pricing for your Plans. Customer Provided Networks can be enabled when creating new Plan by adding following lines to your compose spec file: ``` x-omnistrate-service-plan: features: CUSTOM_NETWORKS: ``` or the following lines on the service spec file: ``` features: CUSTOM_NETWORKS: ``` This feature cannot be modified once Plan is created. ## Configuring Private Networking For each Custom Network created it is possible to configure private network connectivity to allow Customers to use the services with private networking. A way to define private networking is using VPC peering. For more details on how to configure VPC peering you can referent to the [VPC peering guideline](https://docs.omnistrate.com/tenant-management/private-networking/#vpc-peering). Check other options for Private networking in the [Private networking guideline](https://docs.omnistrate.com/tenant-management/private-networking/index.md) # Licensing Protection ## Introduction The Omnistrate Licensing Protection System is designed to ensure that only authorized subscribed users can access and use your software. This system is particularly important for Tenancy Types where customers bring their own accounts (BYOA/BYOC) or run the software on-premises. In these scenarios, software licensing is crucial for protecting intellectual property and enforcing usage policies. ## System Overview The licensing protection system consists of the following key components: - **License File**: A cryptographically signed license file is generated per deployment with a defined expiration. - **Licensing Verification SDK**: An SDK provided by Omnistrate that contains the logic necessary to validate leveraging the Public PKI infrastructure. - **Licensing Renewal Logic**: Automation that will periodically rotate the license while the subscription is active. - **Suspend Subscription Logic**: A management API and UI that can be used to suspend the rotation of the license file, and automation that attempt to retract the license if the deployment or subscription are deleted. ## License Generation and Validation Workflow ### License Feature Configuration The licensing feature needs to be enabled for the Plan. By default the license will be generated with a 7 days expiration. The expiration period means the time it will take for the license to expire if the subscription is suspended or the customer attempts to abuse the system. The license is automatically renewed periodically while the subscription is valid. To ensure the license is unique for your organization, we will use the organization ID and the Product Plan ID to uniquely identify which product the license is generated for. These values need to used to validate the license using the [Licensing SDKs](#sdks). The product plan ID can optionally defined by a value provided in the configuration. Using different product plan keys adds an additional layer of protection that prevents transfer and unauthorized use of licenses across multiple Plans or product you maybe offering to your customers. #### Docker Compose Example ``` x-customer-integrations: licensing: # optional - defaults to 7 days licenseExpirationInDays: 7 # optional - identifier used to add extra security on validation - defaults to product tier id productPlanUniqueIdentifier: '[product plan unique id]' ``` When used on a compose-spec, Omnistrate orchestrates the deployment of the containers and also takes care of mounting the secret and setting the environment variables for verification. #### Service Spec Configuration ``` features: CUSTOMER: licensing: # optional - defaults to 7 days licenseExpirationInDays: 7 # optional - identifier used to add extra security on validation - defaults to product tier id productPlanUniqueIdentifier: '[product plan unique id]' ``` When using Helm or Operator, the secret `service-plan-subscription-license` generated with the license needs to be mounted on `/var/subscription/` in the license feature configuration. ### License Request & Generation When a customer creates a deployment, Omnistrate generates a license file a fixed expiration date. The file is generated uniquely for an Deployment, Subscription and Organization. The license is signed. The private key used for signing the license is from a certificate that is part of the PKI infrastructure. That certificate is generated from [Let's encrypt](https://letsencrypt.org/) to the specific domain `licensing.omnistrate.cloud`. The private key used to sing the license will be rotate periodically. ### License File Distribution & Mounting The license file is distributed along with the public certificate that can be used to validate the signature. The license files are securely delivered to the user and must be placed in a predefined location (e.g., mounted within a container from a secret in the namespaces of the deployment). For custom resources (e.g., Helm), this file needs to be mounted on the data plane to ensure it is always accessible during software execution. Omnistrate SDKs assume the secret is mounted under `/var/subscription/`. The license and certificate files are rotated periodically by Omnistrate while the subscription is valid. Once the subscription is suspended or the deployment deleted the rotation will be paused or the file will be retracted. ### License Validation The public key for the certificate used to sign the license is made available along with the license. The certificate is provided as a secret and has a 90-day expiration. The CA and Intermediary certs, [chain of trust for Let's Encrypt](https://letsencrypt.org/certificates/), that are required to validate that the certificate is valid are included as constants in the SDK, to the validation does not require the CA certs to be updated on the base container image of the service. Once the certificate is proven to be valid, the next step is to check for the signature on the license file. If the signature is valid we can check the license expiration. The SDK will report the license is invalid if the expiration period has passed. If the license is not expired, the SDK will check that the Organization ID and the Product Plan ID (required when the ValidateLicense method is invoked in the SDK) correspond to the values defined in the license. Additionally, the SDK will try to obtain the autogenerated Deployment Instance ID from an environment variable `INSTANCE_ID` and compare that with the value in the license file. This check is performed automatically when using Container image setup, but will be required to inject the Environment variable for Helm, Kustomize and Operators. This is a weak additional check, given that the value is injected and could be tampered, but will discourage license sharing. An SDK is provided in major programming languages to simplify the validation process. For Go, you can use the [Omnistrate Licensing SDK for Go](https://github.com/omnistrate-oss/omnistrate-licensing-sdk-go). For Java, you can use the [Omnistrate Licensing SDK for Java](https://github.com/omnistrate-oss/omnistrate-licensing-sdk-java). Using the SDK you can, with a single call, ensure that the license is valid and protect your software. ### Runtime License Verification On every startup or at predefined intervals, your software can use the SDK to validate the license. Note It is recommended to set up a periodic check on the license file to be able to pause the service whenever the license expires, without the need to wait for a process restart. Examples on how to use the SDK are provided on the SDK code and also we provide some example projects to illustrate how the system works. ### Enforcement Actions If the license check fails (e.g., expired, modified, or missing), your software can take actions such as: - Running in a restricted mode. - Displaying a license violation warning. - Shutting down or preventing further usage. ## Structure of Licensing File The licensing file is a JSON file that contains the following fields: - **ID**: Unique identifier for the license. - **Creation Time**: Timestamp when the license was created. - **Expiration Time**: Timestamp when the license will expire. - **Description**: Description of the Product plan as configured in Omnistrate. - **Instance ID**: Identifier for the deployment associated with the license. - **Organization ID**: Globally unique identifier for your Organization - **Subscription ID**: Identifier for the subscription associated with the license. - **Product Plan Unique ID**: Product Plan Unique Identifier that can be configured. Default to the Product Plan ID (globally unique ID autogenerated by Omnistrate). - **Version**: Version number of the license. The file is signed so any content change will invalidate the signature check. ## SDKs To simplify the license validation process, Omnistrate provides SDKs in major programming languages: - [Omnistrate Licensing SDK for Go](https://github.com/omnistrate-oss/omnistrate-licensing-sdk-go/) - [Omnistrate Licensing SDK for Java](https://github.com/omnistrate-oss/omnistrate-licensing-sdk-java/) ## Examples For practical examples of how to implement the licensing system, you can refer to the following repository: - [Licensing Example for Go](https://github.com/omnistrate-community/licensing-example-go/) - [Licensing Example for Java](https://github.com/omnistrate-community/licensing-example-java/) ## FAQ ### Failure Scenarios This section outlines potential failure scenarios and how Omnistrate handles them to ensure the integrity and availability of the licensing system. #### 1. Rotation is Delayed Omnistrate will make the best effort to retract or rotate licensing files, triggering Alerts to the SaaS Provider in case this is not possible. To prevent issues with rotation we allow a minimum license expiration of 7 days. #### 2. Failed to Retract License File Upon Suspension In case of suspension, if the customer continues operating the software, the license file will expire and the enforcement action will prevent abuse. #### 3. End customer decides to permanently revoke access Dataplane disconnection will not cause unavailability as the license enforcement can be done offline using Omnistrate SDK, but the license file will not be rotated and the license will expire. ### Malicious Actors This section addresses potential threats and questions related to the misuse of the licensing system by malicious actors. #### 1. Can Anyone Create a New License? Omnistrate uses asymmetric encryption to sign the licensing file. Only with access to the private key can new license files be generated that will be able to pass the enforcement actions. The private key used for signing will be from a certificate that will be rotated periodically. That certificate will be part of the Public Key Infrastructure. #### 2. Can Someone Use the License After It Is Revoked? Omnistrate will make the best effort to retract a license for a suspended subscription. In case a malicious actor blocks access to Omnistrate, the license will expire and will become obsolete on the expiration date. #### 3. Can Someone Copy the License and Use It in an External System? With knowledge on what to look into, a malicious actor could copy the license and use it until the expiration date in an external environment. Once the expiration date is reached, the enforcement actions will consider the license obsolete. #### 4. Can someone use the license on a different product? The Organization ID and Product Plan Unique Identifier ensures that the license is being used for the correct Plan and prevents unauthorized usage of the license in different environments or for different products. # Overview Omnistrate offer turn key solution to many capabilities that you can enable to any or all of your Resources. Here are some of the capabilities that we offer: - Reverse proxy to enable https endpoint offload and add TLS termination proxy for secure communication over the internet - MultiZone to place nodes in different availability zones - Autoscaling to enable autoscaling using custom metrics. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/serverless/index.md) - Serverless to make your Resource serverless with seamless scale down to zero. In this mode, we will automatically make your Resource serverless such that if its not in use, we will bring the infrastructure down to zero and bring it back when in use. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/serverless/index.md) - IPWhitelisting to whitelist incoming IPs for incoming traffic. This feature is only available in the enterprise plan - EnableStableEgressIP to provide a stable egress IP for outbound traffic. This feature is currently only available in AWS - ProcessCoreDump to store process core dump in a specific location to debug process crashes - Service account policies for your application to securely talk to cloud-native services. There are multiple permission configurations that Omnistrate provides out of the box: - AWS specific: - MSK_CONNECT - Enables AWS MSK Connect - SECRETS_MANAGER - Enables AWS Secret Manager access - LAMBDA - Enables AWS Lambda access (including permissions necessary to use "Serverless" framework) - SQS - Enables SQS access - GCP specific: - WORKLOAD_IDENTITY_IAM_BINDING - Binds Resources workload identity to IAM service account, granting dataplane additional GCP permissions (such as Logs, Metrics and Secrets) - Service discovery to find and communicate with each different components seamlessly - Custom (Cost) tagging to organize and manage their cloud resources efficiently. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/custom-tagging/index.md) - Backups and Point-in-time restore to save your data periodically and restore it when needed. Note this capability only applies to external Resources. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/pitr/index.md) - Customer networks to allow your Customers to define network partitioning on a dedicated stack. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md) - Custom sidecars to enhance your Product's functionality without changing its core service. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/custom-sidecars/index.md) - Licensing Protection to ensure that only authorized subscribed users can access and use your software. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/licensing-protection/index.md) - Cloud Provider Quotas to manage your cloud quotas. To learn more, please see [here](https://docs.omnistrate.com/runtime-guides/cloud-provider-quotas/index.md) Here is an example configuration: ``` x-omnistrate-capabilities: httpReverseProxy: targetPort: 80 enableMultiZone: true enableStableEgressIP: true autoscaling: maxReplicas: 1 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 scalingMetric: metricEndpoint: "http://localhost:9187/metrics" metricLabelName: "application_name" metricLabelValue: "psql" metricName: "pg_stat_activity_count" serverlessConfiguration: enableAutoStop: true minimumNodesInPool: 5 targetPort: 3306 processCoreDump: /var/lib/data/cores/core.%e.%p.%t backupConfiguration: backupRetentionInDays: 7 backupPeriodInHours: 2 serviceAccountPolicies: aws: - MSK_CONNECT - SECRETS_MANAGER - LAMBDA - SQS gcp: - WORKLOAD_IDENTITY_IAM_BINDING ``` If you don't see your favorite capability above, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com). We would love to understand your use-case and prioritize the support. # Backup and Point-in-time Restore Omnistrate offers a backup and point-in-time restore capability for your Resources. It allows you to recover your end customer Resources to the Recovery Time Objective (RTO) in case of any data loss/corruption or service misconfiguration. You can enable this capability with a few clicks. Once enabled, Omnistrate will automatically and periodically back up your customer's Resources, storing these backups in the same account where the Resources are hosted. If you need an on-demand, persistent recovery artifact for disaster recovery (including cross-region disaster recovery), create a manual deployment snapshot. For details, see [Deployment Snapshots](https://docs.omnistrate.com/operate-guides/deployment-snapshots/index.md). ### Backup Omnistrate enables continuous backup chains for your Resources, ensuring comprehensive backups of all elements, including service data, data plane and control plane configurations. The backups are performed incrementally, using a sliding window approach to optimize storage costs. Additionally, Omnistrate allows you to configure backup retention and Recovery Point Objective (RPO) to best fit your needs. Let's look at above example, if an instance is created on Day 1 at 4:00, Omnistrate begins monitoring your service and initiating the backup chain. In this scenario, continuous backups are taken every hour. The backup chain will keep moving, with the oldest backups being discarded once they reach the maximum retention period. You or your customers can recover Resources to any RPO within the backup chain period. This feature ensures the protection of customer services and data from unexpected interruptions. This capability is only available for Tenant-Aware Resources. To enable backup capability on Tenant-Aware components, add `x-omnistrate-capabilities` in your compose specification. Here is an example: ``` services: app: x-omnistrate-capabilities: backupConfiguration: backupRetentionInDays: 7 backupPeriodInHours: 2 ``` - `backupRetentionInDays`: Number of days to retain the backup - `backupPeriodInHours`: Period in hours to take the backup The backup will be taken on selected Resource and all its dependent Resources. Warning When the backup retention period is updated in a Plan, the change only applies to new backups created after the update. Existing backups will retain their original retention period. In addition, you can enable or disable backups for specific storage volumes. By default, all storage volumes are backed up. To disable backups for a particular storage volume, you can add `disableBackup: true` in the volume configuration. Here is an example using compose spec: ``` services: app: volumes: - source: mysql_master_data target: /var/lib/postgresql/data type: volume x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP2 instanceStorageSizeGi: 200 instanceStorageIOPS: 6000 instanceStorageThroughputMiBps: 800 gcp: instanceStorageType: GCP::PD_SSD instanceStorageSizeGi: 200 azure: instanceStorageType: AZURE::PREMIUM_SSD instanceStorageSizeGi: 200 ``` Note Backup only applies to AWS EBS/GCP/Azure Persistent Disk storage. If you are using other storage types, and would to have backup enabled, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) ### Point-in-time Restore (PITR) Omnistrate allows you to recover your Resources to any RPO within the backup chain period. It restores all components, including service data, configurations, and infrastructure setups, to the exact same running state at the chosen time. This enables you to resume your services immediately, without any additional effort. Typically, the Recovery Time Objective (RTO) is within a couple of minutes and does not impact the source data plane. You or your customers can initiate the restoration by simply selecting the source instance and the target RPO. The restoration process applies not only to the selected Resource but also to all its dependent components. Omnistrate will restore the Resources to the state closest to the selected time, based on the chosen RPO. The restored instance operates independently and is not impacted by the source instance. Note If the source instance is deleted, the automatic backups will also be removed, and you won't be able to restore the Resources. # Serverless ## Overview Omnistrate enables your Resources to have advanced auto-scaling capabilities and scale down to zero with just a few clicks. This allows you to save on infrastructure costs when your customers are not actively using the underlying infrastructure. Additionally, Omnistrate allows you to extend your Resources to behave completely serverless, including preserving any session state. This ensures that your client connections remain unaware when you automatically pause and resume your infrastructure. ## How serverless works Omnistrate's serverless implementation uses an intelligent proxy layer to enable seamless scale-to-zero functionality for your Resources. Here's how it works: ### Architecture overview When you enable serverless for a Resource, Omnistrate automatically deploys a lightweight proxy in front of your application. This proxy serves as the stable endpoint for client connections while your backend Resources can scale down to zero during periods of inactivity. ``` Client → Proxy (Always Running) → Backend Resource (Scales 0-N) ``` ### Scale down to zero process 1. **Monitoring**: The proxy continuously monitors activity on the target port(s) of your Resource 1. **Idle detection**: When no active connections are detected for the configured period, the proxy signals Omnistrate to scale down the backend 1. **Graceful shutdown**: Your Resource instances are gracefully stopped, reducing infrastructure costs to zero 1. **Proxy standby**: The proxy remains active, listening for new connection attempts ### Wake up process 1. **Connection attempt**: A client attempts to connect to your Resource through the proxy endpoint 1. **Automatic provisioning**: The proxy detects the incoming connection and signals Omnistrate to provision backend Resources 1. **Warm pool (optional)**: If configured, a pre-warmed instance from the pool is immediately allocated, reducing startup time from minutes to seconds 1. **Connection establishment**: Once the backend is ready, the proxy establishes the connection and forwards traffic 1. **Transparent to clients**: The client experiences a brief connection delay but remains unaware of the underlying scale-to-zero operation ### Benefits of the proxy architecture - **Stable endpoints**: Your clients always connect to the same proxy endpoint, regardless of backend state - **Seamless scaling**: Backend Resources can scale up and down without disrupting client connectivity - **Cost optimization**: Only pay for compute resources when actively in use - **Session preservation (Advanced mode)**: The proxy can maintain session state during scale-down operations ## Scale down to zero Autoscaling requires you to have a minimum of one machine, but if you want to scale down to zero and automatically restart machines during load, you can enable the serverless capability for the corresponding Resource instances. Enabling `serverlessConfiguration` will enable auto-stop and auto-wakeup capabilities for your Resource with just one click. There are two serverless modes: Basic and Advanced. ## Configuration You enable basic serverless by attaching `serverlessConfiguration` to a Resource under `x-omnistrate-capabilities`. Even though the configuration lives on a Resource, basic serverless scales the entire deployment down to zero and wakes the entire deployment back up. In most cases, you should configure it on a single Resource that best represents the deployment's incoming traffic rather than copying it to every service. Here is an example using compose spec: ``` services: app: x-omnistrate-capabilities: serverlessConfiguration: enableAutoStop: true targetPort: 3306 ``` Note Serverless capability is available for Hosted SaaS and BYOC deployments. It is not available for Air-gapped deployments. Basic serverless scope Basic serverless is a deployment-wide behavior. You usually configure it on one representative Resource and port that can act as the wake-up signal for the full deployment. Custom DNS and serverless Custom DNS aliases do not currently trigger basic serverless wake-up. If you combine custom DNS with serverless, use the generated Omnistrate endpoint to wake the deployment. Do not assume requests to the custom DNS alias will resume a scaled-to-zero deployment. The default serverless mode only works for TCP-based services and uses the number of active connections to a target port as a metric to scale down to zero. For custom metrics, see [Advanced serverless](#advanced-serverless) for more details. ## Testing serverless functionality Create an instance of the serverless Resource. Once it's up and running, check the connectivity. Wait for the configured period of inactivity for the infrastructure to be automatically stopped. You can confirm this by checking the state of the Resource instance. As soon as new activity resumes, the underlying Resource instance should start automatically within several seconds. ## Faster wakeup using warm pool To speed up the wakeup time, Omnistrate offers a warm pool capability that you can enable for faster resume times. ``` minimumNodesInPool: 5 ``` - `minimumNodesInPool`: Configure a static number of machines in the pool. You can change this value at any time programmatically through APIs or manually through the UI. ## Demo Example Here is a demo example using PostgreSQL: ### Achieving full elasticity Scale down to zero enables your underlying software to go from 0 to 1 when in use and back to 0 when not in use. However, in practice, you may want to scale from 0 to N as the load increases and back to 0 when not in use. To achieve this, enable auto-scaling in addition to scale down to zero. Here is an example configuration using compose spec: ``` services: app: x-omnistrate-capabilities: autoscaling: maxReplicas: 5 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 scalingMetric: metricEndpoint: "http://localhost:9187/metrics" metricLabelName: "application_name" metricLabelValue: "custom_metric_label" metricName: "custom_metric_name" serverlessConfiguration: targetPort: 3306 enableAutoStop: true minimumNodesInPool: 5 ``` In this pattern: - `autoscaling` controls scaling between 1 and `maxReplicas` while the deployment is running - `serverlessConfiguration` controls the transition between zero and a running deployment - the serverless proxy decides when to wake the deployment based on traffic to the configured `targetPort` ## Advanced serverless You may need advanced serverless if you require any of the following capabilities: - **Resume client state after scaling to zero**: For example, if you want to scale down the infrastructure for your database application on idle connections but have prepared statements or temporary tables as part of the session state, the default behavior will lose all connection state and client connections will fail. With advanced serverless, you can store the connection state so that on resume, clients continue to behave the same way, making them completely unaware of the scale down to zero. - **Custom metrics**: Use custom metrics to determine when to scale down to zero or resume. - **Proxy configuration control**: Take control of the proxy configuration, including its image, infrastructure, capabilities, and managing the lifecycle of the proxy Resource. To enable advanced serverless using compose spec: ``` services: app: serverlessConfiguration: referenceProxyKey: "proxyResourceKey" portsMappingProxyConfig: numberOfPortsPerCluster: 4 maxNumberOfClustersPerProxyInstance: 1 proxyStorageConfig: aws: storageType: "s3" ``` - `referenceProxyKey`: Reference to the proxy Resource that will be used to listen to client traffic - `proxyStorageConfig`: State store to keep the connection state(s) - `portsMappingProxyConfig`: Configurations to map proxy Resource ports to the database application Info Advanced serverless mode is only available in the enterprise plan. Please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) to learn more. In summary, if you require any customization, you will likely need advanced mode. For simple use cases, the basic serverless mode should suffice. # Tenant Management # Build your own customer portal ## Building a customer portal from scratch You can build a completely custom Customer Portal using the Omnistrate REST APIs. This approach gives you full control over the user experience, branding, and functionality of your customer-facing portal. #### Benefits of building from scratch - **Complete Customization**: Full control over UI/UX design and user workflows - **Brand Integration**: Seamlessly integrate with your existing brand guidelines and design system - **Advanced Features**: Implement complex business logic and custom features specific to your needs - **Technology Choice**: Use any frontend framework, programming language, or technology stack - **Data Control**: Complete ownership of customer data and portal analytics #### Prerequisites Before building your Customer Portal, ensure you have: - An active Omnistrate account with at least one SaaS Product defined - A production environment released for your service - Basic understanding of REST API integration - Familiarity with your chosen frontend technology stack #### Development approach This guide demonstrates portal development using REST API calls, which you can implement in any programming language or framework. The examples use command-line tools to illustrate the API interactions, but the same concepts apply when building web applications, mobile apps, or other portal interfaces. #### Recommended tools Restish We will use [Restish](https://rest.sh/#/) to walk through the APIs and design a basic Customer Portal for your SaaS. Restish is a CLI tool that makes it easy to interact with REST APIs, allowing us to demonstrate the complete customer journey from sign-up to resource deployment without writing frontend code. Why Restish for This Guide? - **Rapid Prototyping**: Quickly test API workflows without building a full application - **Clear Examples**: Each API call is explicit and easy to understand - **Transferable Knowledge**: The same API patterns work in any programming language - **Configuration Management**: Easy to manage multiple API profiles and authentication tokens ## Configure your customer portal APIs The Omnistrate platform provides 2 separate set of API models for your use: - **APIs to operate your service** - **Impersonation APIs to operate the Customer Portal on behalf of your customers** Let's setup two separate API profiles in Restish for these use-cases. 1. **APIs to operate your service** ``` restish api configure service-operator https://api.omnistrate.cloud/2022-09-01-00/fleet ``` 1. **Impersonation APIs to operate the Customer Portal on behalf of your customers** ``` restish api configure portal https://api.omnistrate.cloud/2022-09-01-00 ``` The default profile is for you as a service owner to onboard (sign-up, validate token, etc.) a customer. The rest of the APIs in this model are operated in the context of the customer. ## Set up authentication Next, we will setup the authentication mechanism for each of these APIs. Add your username / password to a JSON file `service-signin.json`: ``` { "email": "your-email", "password": "your-password" } ``` Sign-in as the service owner: ``` restish portal signin-api-signin " } ``` ______________________________________________________________________ Now, configure both the APIs with the JWT token: ``` restish api edit # Modify the JSON file to use the JWT token { "$schema": "https://rest.sh/schemas/apis.json", "portal": { "base": "https://api.omnistrate.cloud/2022-09-01-00", "profiles": { "default": { "headers": { "Authorization": "Bearer " } } } }, "service-operator": { "base": "https://api.omnistrate.cloud/2022-09-01-00/fleet", "profiles": { "default": { "headers": { "Authorization": "Bearer " } } } } } ``` ______________________________________________________________________ Verify that the APIs are setup correctly by checking the caller identity: ``` restish portal users-api-describe-user HTTP/2.0 200 OK Content-Length: 849 Content-Security-Policy: default-src 'none' Content-Type: application/json Date: Mon, 07 Oct 2024 06:02:14 GMT { createdAt: "2023-04-13T10:07:33Z" email: "demo@omnistrate.com" id: "user-18ZZD6JBrS" lastModifiedAt: "2024-10-04T21:30:15Z" name: "Omnistrate Demo" orgDescription: "OmnistrateDemo" orgName: "OmnistrateDemo" orgURL: "https://www.omnistrate.com" planName: "ENTERPRISE" roleType: "root" } ``` ## Onboard your first customer Now that the APIs are setup, let's onboard your first customer. Omnistrate provides an event loop to manage the lifecycle of your customers. Every customer activity (sign-up, invite, subscription request, etc.) generates an event that you can poll in your app and take appropriate action. Use the `portal` APIs to create a new customer. Add the following JSON payload to a file `customer-signup.json`: ``` { "companyDescription": "The best eCommerce website", "companyUrl": "https://www.example.com", "email": "admin@example.com", "legalCompanyName": "eCommerce", "name": "John Doe", "password": "" } ``` ``` restish portal users-api-customer-signup " user_email: "admin@example.com" user_id: "user-dcvEakUjz0" user_name: "" user_organization: "" } eventType: "CustomerSignUp" orgID: "org-4aoZZqPP3T" orgName: "Omnistrate" orgURL: "" priority: "" time: "2024-10-07T06:10:38Z" userEmail: "admin@example.com" userID: "user-dcvEakUjz0" userName: "" } ] } ``` ______________________________________________________________________ Your service can send an email at this point to the customer with the validation token to complete the sign-up process. For now, let's explicitly validate the customer with the token in the event payload. Add the following JSON payload to a file `customer-validate.json`: ``` { "email": "admin@example.com", "token": "" } ``` ``` restish portal signup-api-validate-token " } ``` ``` restish portal users-api-customer-signin " } }, "default": { "headers": { "authorization": "Bearer " } } } } ... ``` ______________________________________________________________________ Impersonate the customer and fetch the details of the customer: ``` restish portal users-api-describe-user --rsh-profile customer HTTP/2.0 200 OK Content-Length: 420 Content-Security-Policy: default-src 'none' Content-Type: application/json Date: Mon, 07 Oct 2024 06:32:44 GMT { createdAt: "2024-10-07T06:10:38Z" email: "admin@example.com" id: "user-dcvEakUjz0" lastModifiedAt: "2024-10-07T06:16:29Z" name: "John Doe" orgDescription: "The best eCommerce website" orgFavIconURL: "" orgId: "org-z0sasy6jke" orgLogoURL: "" orgName: "eCommerce" orgPrivacyPolicy: "" orgSupportEmail: "" orgTermsOfUse: "" orgURL: "https://www.example.com" planName: "STARTER_NO_COMMIT" roleType: "root" } ``` ## Manage users for your customers Omnistrate provides the concept of an organization for your customers to manage their users. Once onboarded, your customers can invite other team members to join their organization on Omnistrate through the Customer Portal APIs. This provides a unified view for you and your customers to manage the billing, usage, and access control of the users of your service. #### Invite a customer Let's impersonate the customer we just onboarded and invite another user to join their organization. Add the following JSON payload to a file `customer-invite.json`: ``` { "email": "user1@example.com" } ``` ``` restish portal users-api-customer-invite-user " } ``` ``` restish portal users-api-customer-signup " user_email: "user1@example.com" user_id: "user-qUClcuYyBW" user_name: "" user_organization: "" } eventType: "CustomerSignUp" orgID: "org-4aoZZqPP3T" orgName: "Omnistrate" orgURL: "" priority: "" time: "2024-10-07T06:43:23Z" userEmail: "user1@example.com" userID: "user-qUClcuYyBW" userName: "" } ] } ``` ______________________________________________________________________ Validate the invited user. Add the following JSON payload to a file `customer-invited-validate.json`: ``` { "email": "user1@example.com", "token": "" } ``` ``` restish portal signup-api-validate-token ", "postgresqlUsername": "" } } ``` ______________________________________________________________________ Let's fetch the details of the deployment using the deployment / resource-instance ID: ``` restish portal resource-instance-api-describe-resource-instance --subscription-id sub-ImCvoojwpm sp-JGcqr1hPxX postgres-saas v1 production postgres-saas-omnistrate-hosted postgres-saas-postgres-saas-omnistrate-hosted-model-omnistrate-dedicated-tenancy cluster instance-5i4m1cwrp --rsh-profile customer HTTP/2.0 200 OK Content-Security-Policy: default-src 'none' Content-Type: application/json Date: Mon, 07 Oct 2024 07:28:20 GMT { active: true cloud_provider: "aws" createdByUserId: "user-dcvEakUjz0" createdByUserName: "John Doe" created_at: "2024-10-07T07:24:08Z" detailedNetworkTopology: { r-DbJa2uX0L6: { allowedIPRanges: ["0.0.0.0/0"] clusterEndpoint: "pgadmin.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" clusterPorts: [80, 443] customDNSEndpoint: { enabled: false } hasCompute: true main: false networkingType: "PUBLIC" nodes: [ { availabilityZone: "us-east-2a" detailedHealth: { ConnectivityStatus: "HEALTHY" DiskHealth: "HEALTHY" LoadHealth: "POD_NORMAL" NodeHealth: "HEALTHY" ProcessHealth: "HEALTHY" ProcessLiveness: "HEALTHY" } endpoint: "pgadmin-0.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" healthStatus: "HEALTHY" id: "pgadmin-0" ports: [80] status: "RUNNING" storageSize: 30 } ] privateNetworkCIDR: "172.20.0.0/16" publiclyAccessible: true resourceInstanceMetadata: null resourceKey: "pGAdmin" resourceName: "PGAdmin" } r-M7LY75veyE: { allowedIPRanges: [] clusterEndpoint: "" customDNSEndpoint: { enabled: false } hasCompute: false main: true networkingType: "" privateNetworkCIDR: "172.20.0.0/16" publiclyAccessible: true resourceInstanceMetadata: null resourceKey: "cluster" resourceName: "Cluster" } r-WyuLUOnWVy: { allowedIPRanges: ["0.0.0.0/0"] clusterEndpoint: "r-wyuluonwvy.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" clusterPorts: [5432] customDNSEndpoint: { enabled: false } hasCompute: true main: false networkingType: "PUBLIC" nodes: [ { availabilityZone: "us-east-2a" detailedHealth: { ConnectivityStatus: "HEALTHY" DiskHealth: "HEALTHY" LoadHealth: "POD_IDLE" NodeHealth: "HEALTHY" ProcessHealth: "HEALTHY" ProcessLiveness: "HEALTHY" } endpoint: "writer-0.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" healthStatus: "HEALTHY" id: "writer-0" ports: [5432] status: "RUNNING" storageSize: 30 } ] privateNetworkCIDR: "172.20.0.0/16" publiclyAccessible: true resourceInstanceMetadata: null resourceKey: "writer" resourceName: "Writer" } r-aRgU5ukzZV: { allowedIPRanges: ["0.0.0.0/0"] clusterEndpoint: "r-argu5ukzzv.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" clusterPorts: [5432] customDNSEndpoint: { enabled: false } hasCompute: true main: false networkingType: "PUBLIC" nodes: [ { availabilityZone: "us-east-2a" detailedHealth: { ConnectivityStatus: "HEALTHY" DiskHealth: "HEALTHY" LoadHealth: "POD_IDLE" NodeHealth: "HEALTHY" ProcessHealth: "HEALTHY" ProcessLiveness: "HEALTHY" } endpoint: "reader-0.instance-5i4m1cwrp.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" healthStatus: "HEALTHY" id: "reader-0" ports: [5432] status: "RUNNING" storageSize: 30 } ] privateNetworkCIDR: "172.20.0.0/16" publiclyAccessible: true resourceInstanceMetadata: null resourceKey: "reader" resourceName: "Reader" } r-obsrvpt48sqr0hcce: { allowedIPRanges: ["0.0.0.0/0"] clusterEndpoint: "streamer.hc-pelsk80ph.us-east-2.aws.f2e0a955bb84.cloud" clusterPorts: [8080] customDNSEndpoint: { enabled: false } hasCompute: false main: false networkingType: "PUBLIC" privateNetworkCIDR: "172.20.0.0/16" publiclyAccessible: true resourceInstanceMetadata: null resourceKey: "omnistrateobserv" resourceName: "Omnistrate Observability" } } highAvailability: false id: "instance-5i4m1cwrp" last_modified_at: "2024-10-07T07:28:02Z" network_type: "PUBLIC" productTierFeatures: { LOGS: { enabled: true featureName: "LOGS" label: ["ONLINE"] scope: "CUSTOMER" } METRICS: { enabled: true featureName: "METRICS" label: ["ONLINE"] scope: "CUSTOMER" } } region: "us-east-2" result_params: { dbName: "abc" instanceType: "t4g.small" pgadminEmailAddress: "alok@omnistrate.com" postgresqlUsername: "homepage" } status: "RUNNING" subscriptionId: "sub-ImCvoojwpm" } ``` This provides a detailed view of the Postgres cluster deployed for your customer including the network topology, health status, and the endpoints for the cluster. You can customize the experience for your customer by providing them with the necessary information to connect to the cluster. ## Implement customer RBAC Omnistrate allows more fine-grained RBAC within the context of a customer organization. A customer can invite other users within (or outside) their org to share a subscription with them. This provides an easy way for your customers to pay for your service while providing easy access to their team members. In this example, let's add the user we invited earlier to share the subscription with the customer that owns it. #### Set up the invited user First, let's login as the invited user. Add the following JSON payload to a file `customer-invited-signin.json`: ``` { "email": "user1@example.com", "password": "" } ``` ``` restish portal users-api-customer-signin " } ``` Add the invited user to the Restish profile: ``` restish api edit # Modify the JSON file to use the JWT token { "$schema": "https://rest.sh/schemas/apis.json", "portal": { "base": "https://api.omnistrate.cloud/2022-09-01-00", "profiles": { "customer": { "headers": { "authorization": "Bearer " } }, "default": { "headers": { "authorization": "Bearer " } }, "user1": { "headers": { "authorization": "Bearer " } } } } } ``` ______________________________________________________________________ Let's verify the new user's identity: ``` restish portal users-api-describe-user --rsh-profile user1 HTTP/2.0 200 OK Content-Length: 424 Content-Security-Policy: default-src 'none' Content-Type: application/json Date: Mon, 07 Oct 2024 07:34:52 GMT { createdAt: "2024-10-07T06:40:35Z" email: "user1@example.com" id: "user-qUClcuYyBW" lastModifiedAt: "2024-10-07T06:47:59Z" name: "Jane Smith" orgDescription: "The best eCommerce website" orgFavIconURL: "" orgId: "org-z0sasy6jke" orgLogoURL: "" orgName: "eCommerce" orgPrivacyPolicy: "" orgSupportEmail: "" orgTermsOfUse: "" orgURL: "https://www.example.com" planName: "STARTER_NO_COMMIT" roleType: "member" } ``` #### Share the subscription with the invited user Let's attempt to describe the Postgres cluster as the invited user: ``` restish portal resource-instance-api-describe-resource-instance --subscription-id sub-ImCvoojwpm sp-JGcqr1hPxX postgres-saas v1 production postgres-saas-omnistrate-hosted postgres-saas-postgres-saas-omnistrate-hosted-model-omnistrate-dedicated-tenancy cluster instance-5i4m1cwrp --rsh-profile user1 HTTP/2.0 403 Forbidden Content-Length: 157 Content-Security-Policy: default-src 'none' Content-Type: application/json Date: Mon, 07 Oct 2024 07:36:13 GMT { fault: false id: "" message: "unauthorized: not authorized to perform DESCRIBE action on instance" name: "forbidden" temporary: false timeout: false } ``` As expected, the invited user does not have the necessary permissions to describe the Postgres cluster. ______________________________________________________________________ Let's add the user to the subscription `sub-ImCvoojwpm` with **read-only** rights. Add the following JSON payload to a file `customer-invite-user.json`: ``` { "email": "user1@example.com", "roleType": "reader" } ``` ``` restish portal consumption-user-api-invite-user sub-ImCvoojwpm .com`) You can customize your SaaS domain to make it easier for SEO and your tenants to remember. To host your control plane UX on your domain, setup CNAME records with your domain provider. We generate a domain name for each environment that you configure and you can configure each environment separately. Please note that all non-production environments are private and accessible only to your organization. If you want to invite customers to self-serve, you can create a public environment, and we will automatically allow anyone to sign up and start using your application. To configure this you can update the custom domain parameters. Alternatively, if you are hosting [Customer Portal](https://github.com/omnistrate-oss/customer-portal) yourself, please follow [README](https://github.com/omnistrate-oss/customer-portal/blob/master/README.md) - **Sender Email Configuration:** For sending emails to your customers from the Customer Portal. You can use your own SMTP server to configure email sends and ensure all communication with your customers is done using your domain. To configure this you can update the parameters: - SMTP Username for authentication - SMTP Password for authentication - SMTP Host - SMTP Port - From Email: the sender email address for outgoing emails - **Configure Identity Providers** Omnistrate's Customer Portal an be customized to define your own Identity Providers. You can use any OpenID Connect compatible Identity Provider, including Google, Github, AWS Cognito, Microsoft Entra, Auth0, Okta, KeyCloak and others. See more information on how to configure your Identity Providers [here](https://docs.omnistrate.com/tenant-management/identity-providers/index.md) If you want to exclusively use OpenID Connect compatible Identity Provider, you can disable User Name and Password on the Portal Configuration. This will disallow customers to sign up using Omnistrate native Identity Provider. - **Google Analytics Tag ID** Get your own usage analytics customizing the google tag id. ### Use your own portal For complete control over your customer experience, you have multiple options for creating and hosting your own custom portal. Whether you fork our open source repository or build a completely custom solution, Omnistrate can host your portal using custom container images. **Requirements for Custom Images:** - **Public Accessibility**: Your container image must be publicly available so that Omnistrate can pull and deploy it. **Configuration Steps:** 1. Build and publish your custom portal container image to a public registry 1. In your Omnistrate Customer Portal configuration, specify the custom image details 1. **Image Name** (e.g. `ghcr.io/your-org/custom-portal` or `your-org/custom-portal`) 1. **Image Tag** (e.g. `v1.0.0`, `latest`, or specific version tags) This approach gives you the flexibility to create a fully customized customer experience while benefiting from Omnistrate's managed hosting, scaling, and infrastructure capabilities. Note When you provide a Custom Image, you must also specify your Sender Email Configuration to ensure proper email delivery from your custom portal. Warning Ensure your Customer Portal container image is publicly accessible before configuring it in Omnistrate. # Identity providers (single sign-on) Omnistrate allows configuring Single Sign-on (SSO) for your Customer Portal. This feature enables end users to authenticate into the Customer Portal using existing identity provider credentials, such as Google, GitHub, Microsoft Entra, and many others. For enterprise customers, this means they can leverage their existing enterprise identity (such as Microsoft Entra / Active Directory, Okta, Auth0, or Keycloak) to provide seamless authentication. This integration allows you to maintain centralized user management, enforce corporate security policies, and provide a familiar login experience. This feature is an addition to the username/password authentication mechanism provided by default. You can disable the default username/password authentication updating the [Customer Portal configuration](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) and rely solely on identity providers for user authentication. ## What is OpenID Connect? SSO uses the [OpenID Connect protocol](https://openid.net/connect/) for authentication, which is built on top of OAuth 2.0 and provides standardized identity verification. ### How OpenID Connect authentication works When a user clicks on an identity provider login button in your Customer Portal, the following authentication flow occurs: 1. **Initiation**: The user is redirected to the identity provider's authorization endpoint 1. **Authentication**: The user authenticates with their identity provider (e.g., enters Google credentials) 1. **Authorization**: The identity provider asks the user to authorize your application to access their profile information 1. **Callback**: The identity provider redirects back to your User Portal with an authorization code 1. **Token Exchange**: Omnistrate exchanges the authorization code for an access token and ID token at the provider's token endpoint 1. **User Information**: Omnistrate retrieves the user's profile information from the provider's user info endpoint or the token 1. **Session Creation**: A user session (JWT token) is created in your Customer Portal using the verified identity information This flow ensures secure authentication while maintaining user privacy and following industry standards. ## Supported identity providers Omnistrate supports the following identity providers: - **Google** - Google Workspace and personal Google accounts - **GitHub** - GitHub OAuth applications - **Amazon Cognito** - AWS managed identity service - **Microsoft Entra** (formerly Azure AD) - Microsoft's identity platform - **Okta** - Enterprise identity management - **Auth0** - Identity-as-a-Service platform - **Keycloak** - Open-source identity and access management - **OpenID Connect** - Any OpenID Connect compliant provider Once the SSO Identity Provider(s) are configured, users can log in to your SaaS application using the buttons on the login page. ## Configuration settings When configuring an identity provider, you'll need to provide the following information: ### Required settings - **Client ID**: The client ID of the identity provider application - **Client Secret**: The client secret of the identity provider application ### OpenID Connect endpoints The behavior of these endpoints depends on the provider type: **Fixed Endpoint Providers** (Google, GitHub): These providers have well-known, fixed endpoints that Omnistrate automatically configures. You don't need to specify these endpoints manually as they are built into the platform. **Configurable Endpoint Providers** (Amazon Cognito, Microsoft Entra, Okta, Auth0, Keycloak, Generic OpenID Connect): For these providers, you can specify custom endpoints or let Omnistrate auto-discover them from the provider's OpenID Connect discovery document. - **Authorization Endpoint**: OAuth2 endpoint to initiate the authentication request - **Token Endpoint**: Endpoint used to exchange authorization code for access and ID tokens - **User Info Endpoint**: Endpoint to fetch authenticated user's information ### Optional settings These settings allow you to customize how the identity provider appears and behaves in your Customer Portal: - **Scopes**: OAuth scopes (space separated) defining what information will be accessed from the identity provider (e.g., `openid email profile`). This determines what user information Omnistrate can retrieve and display in the Customer Portal. Common scopes include: - `openid`: Required for OpenID Connect authentication flow (enables ID token generation) - `email`: Access to user's email address (displayed in user profile) - `profile`: Access to basic profile information like name and avatar (displayed in user profile and throughout the portal). - **LoginButton Icon URL**: A URL to a custom image used as the icon on the login button in your Customer Portal. This allows you to: - Brand the login experience with your identity provider's logo - Use custom icons that match your portal's design theme - Provide visual recognition for users familiar with the identity provider - If not specified, Omnistrate will use a default icon for the provider type - **Login Button Text**: Custom text displayed on the login button in your Customer Portal. This allows you to: - Customize the call-to-action text (e.g., "Continue with Google" vs "Sign in with Google") - Use language that matches your brand voice - Provide context-specific messaging for different user types - If not specified, Omnistrate will use default text like "Continue with [Provider Name]" - **Email Identifiers**: Configure specific email domains (comma-separated) that should authenticate using this provider (e.g., `omnistrate.com, google.com`). This setting: - Automatically directs users with matching email domains to this identity provider - Helps streamline the login experience for enterprise users - Can be used to enforce corporate authentication policies - Allows different email domains to use different identity providers - If not specified, the identity provider will be available to all users regardless of email domain ## Provider setup instructions For all identity providers, you must create an **OAuth 2.0 Application configured for Web Applications**. This is essential for the OpenID Connect authentication flow to work properly with your Customer Portal. Each provider has specific steps to create and configure these OAuth applications as detailed below. ### Google Google uses fixed, well-known OpenID Connect endpoints that are automatically configured by Omnistrate. You only need to provide the Client ID and Client Secret. 1. Go to the [Google Cloud Console](https://console.cloud.google.com/) 1. Create a new project or select an existing one 1. Enable the Google+ API 1. Go to **Credentials** → **Create Credentials** → **OAuth 2.0 Client IDs** 1. Configure the OAuth consent screen if not already done 1. Set **Application type** to "Web application" 1. Add your SaaS domain URL to **Authorized JavaScript origins** 1. Add your SaaS domain URL with `/idp-auth` endpoint to **Authorized redirect URIs** 1. Copy the **Client ID** and **Client Secret** **Fixed Endpoints** (automatically configured): - Authorization: `https://accounts.google.com/o/oauth2/v2/auth` - Token: `https://oauth2.googleapis.com/token` - User Info: `https://openidconnect.googleapis.com/v1/userinfo` **Required Scopes**: `openid email profile` **Documentation**: [Google OAuth 2.0 Setup](https://developers.google.com/identity/protocols/oauth2?sjid=10696970028599612746-NC) Note If your Google application is set to "Internal" user type, only users within your organization (@your-organization.com) will be able to authenticate with this Identity Provider. Omnistrate is not able to verify your Identity Provider credentials in this case. ### GitHub GitHub uses fixed endpoints that are automatically configured by Omnistrate. You only need to provide the Client ID and Client Secret. 1. Go to your GitHub account **Settings** → **Developer settings** → **OAuth Apps** 1. Click **New OAuth App** 1. Fill in the application details: 1. **Application name**: Your SaaS application name 1. **Homepage URL**: Your SaaS domain URL 1. **Authorization callback URL**: Your SaaS domain URL with `/idp-auth` endpoint 1. Click **Register application** 1. Copy the **Client ID** and generate a **Client Secret** **Fixed Endpoints** (automatically configured): - Authorization: `https://github.com/login/oauth/authorize` - Token: `https://github.com/login/oauth/access_token` - User Info: `https://api.github.com/user` **Required Scopes**: `read:user user:email` **Documentation**: [GitHub OAuth Apps](https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/creating-an-oauth-app) ### Amazon Cognito 1. Go to the [AWS Cognito Console](https://console.aws.amazon.com/cognito/) 1. Select your User Pool or create a new one 1. Go to **App integration** → **App clients** 1. Create a new app client or select an existing one 1. Configure the app client: 1. Enable **Authorization code grant** 1. Add your SaaS domain URL with `/idp-auth` to **Allowed callback URLs** 1. Add required scopes: `openid`, `email`, `profile` 1. Note down the **Client ID** and **Client Secret** 1. Find your Cognito domain under **App integration** → **Domain** **Endpoints**: - Authorization: `https://your-domain.auth.region.amazoncognito.com/oauth2/authorize` - Token: `https://your-domain.auth.region.amazoncognito.com/oauth2/token` - User Info: `https://your-domain.auth.region.amazoncognito.com/oauth2/userInfo` **Documentation**: [Amazon Cognito OAuth 2.0](https://docs.aws.amazon.com/cognito/latest/developerguide/cognito-user-pools-app-integration.html) ### Microsoft Entra (Azure AD) 1. Go to the [Azure Portal](https://portal.azure.com/) 1. Navigate to **Azure Active Directory** → **App registrations** 1. Click **New registration** 1. Configure the application: 1. **Name**: Your SaaS application name 1. **Supported account types**: Choose appropriate option 1. **Redirect URI**: Web - Your SaaS domain URL with `/idp-auth` 1. Go to **Certificates & secrets** → **Client secrets** → **New client secret** 1. Copy the **Application (client) ID** and **Client secret value** **Endpoints** (replace `{tenant-id}` with your tenant ID): - Authorization: `https://login.microsoftonline.com/{tenant-id}/oauth2/v2.0/authorize` - Token: `https://login.microsoftonline.com/{tenant-id}/oauth2/v2.0/token` - User Info: `https://graph.microsoft.com/oidc/userinfo` **Required Scopes**: `openid email profile` **Documentation**: [Microsoft Identity Platform](https://docs.microsoft.com/en-us/azure/active-directory/develop/quickstart-register-app) ### Okta 1. Go to your Okta Admin Console 1. Navigate to **Applications** → **Applications** 1. Click **Create App Integration** 1. Select **OIDC - OpenID Connect** and **Web Application** 1. Configure the application: 1. **App integration name**: Your SaaS application name 1. **Grant type**: Authorization Code 1. **Sign-in redirect URIs**: Your SaaS domain URL with `/idp-auth` 1. Copy the **Client ID** and **Client secret** **Endpoints** (replace `{your-okta-domain}` with your Okta domain): - Authorization: `https://{your-okta-domain}/oauth2/default/v1/authorize` - Token: `https://{your-okta-domain}/oauth2/default/v1/token` - User Info: `https://{your-okta-domain}/oauth2/default/v1/userinfo` **Required Scopes**: `openid email profile` **Documentation**: [Okta OIDC Web App](https://developer.okta.com/docs/guides/sign-into-web-app-redirect/asp-net-core-3/main/) ### Auth0 1. Go to your [Auth0 Dashboard](https://manage.auth0.com/) 1. Navigate to **Applications** → **Applications** 1. Click **Create Application** 1. Choose **Regular Web Applications** 1. Go to the **Settings** tab: 1. **Allowed Callback URLs**: Your SaaS domain URL with `/idp-auth` 1. **Allowed Web Origins**: Your SaaS domain URL 1. Copy the **Client ID** and **Client Secret** **Endpoints** (replace `{your-auth0-domain}` with your Auth0 domain): - Authorization: `https://{your-auth0-domain}/authorize` - Token: `https://{your-auth0-domain}/oauth/token` - User Info: `https://{your-auth0-domain}/userinfo` **Required Scopes**: `openid email profile` **Documentation**: [Auth0 Regular Web App](https://auth0.com/docs/get-started/auth0-overview/create-applications/regular-web-apps) ### Keycloak 1. Access your Keycloak Admin Console 1. Select your realm or create a new one 1. Go to **Clients** → **Create client** 1. Configure the client: 1. **Client type**: OpenID Connect 1. **Client ID**: Your application identifier 1. In the client settings: 1. **Access Type**: Confidential 1. **Valid Redirect URIs**: Your SaaS domain URL with `/idp-auth` 1. **Web Origins**: Your SaaS domain URL 1. Go to **Credentials** tab and copy the **Secret** **Endpoints** (replace `{keycloak-server}` and `{realm}` with your values): - Authorization: `https://{keycloak-server}/auth/realms/{realm}/protocol/openid-connect/auth` - Token: `https://{keycloak-server}/auth/realms/{realm}/protocol/openid-connect/token` - User Info: `https://{keycloak-server}/auth/realms/{realm}/protocol/openid-connect/userinfo` **Required Scopes**: `openid email profile` **Documentation**: [Keycloak OpenID Connect](https://www.keycloak.org/securing-apps/oidc-layers) ### Generic OpenID Connect For any OpenID Connect compliant provider: 1. Register your application with the identity provider 1. Configure the redirect URI to your SaaS domain URL with `/idp-auth` 1. Obtain the **Client ID** and **Client Secret** 1. Find the OpenID Connect discovery document (usually at `/.well-known/openid_configuration`) 1. Extract the required endpoints from the discovery document **Required Information**: - Client ID and Client Secret - Authorization Endpoint - Token Endpoint - User Info Endpoint - Supported scopes (minimum: `openid email`, recommended: `openid email profile`) ## Setting up identity providers in Omnistrate 1. Navigate to the Identity Providers section in [Omnistrate](https://omnistrate.cloud) 1. Click **Add Identity Provider** 1. Select your identity provider from the dropdown 1. Fill in the required configuration: 1. **Client ID**: From your identity provider application 1. **Client Secret**: From your identity provider application 1. **Scopes**: Required scopes (typically `openid email profile`) 1. **Endpoints**: Auto-filled for known providers, manual entry for generic OpenID Connect 1. Configure optional settings: 1. **Login Button Text**: Custom text for the login button 1. **Login Button Icon URL**: Custom icon for the login button 1. **Email Identifiers**: Restrict to specific email domains 1. Click **Verify** to test the configuration 1. Click **Save** to enable the identity provider ## Per-environment identity provider configuration Omnistrate supports configuring different identity providers (IdPs) for each environment (for example, development, staging, and production). This allows you to: - Use a **test identity provider** in development environments for easier debugging and faster iteration. - Enforce **corporate SSO** only in production, while allowing password-based login in development. - Configure **separate client credentials** per environment to isolate authentication traffic. ### How it works Each environment in Omnistrate maintains its own set of identity provider configurations. When you add or modify an identity provider, the configuration applies only to the environment you are currently working in. ### Setting up per-environment identity providers 1. Navigate to the Identity Providers section in [Omnistrate](https://omnistrate.cloud). 1. Select the **environment** you want to configure from the environment dropdown in the top navigation bar. 1. Add or modify identity providers for that environment following the standard [setup instructions](#setting-up-identity-providers-in-omnistrate). 1. Repeat for each environment as needed. Tip Use a dedicated OAuth application per environment in your identity provider (for example, separate Google OAuth apps for development and production). This ensures callback URLs, usage metrics, and access controls are isolated between environments. Note Changes to identity provider configuration in one environment do not affect other environments. Each environment operates independently. ## Advanced security features By integrating with enterprise identity providers, you can bring advanced security features to your Customer Portal: ### SAML Support Many identity providers (such as Microsoft Entra, Okta, Auth0, and Keycloak) support SAML (Security Assertion Markup Language) authentication. When users authenticate through these providers, they benefit from SAML's enterprise-grade security features, including: - **Single Sign-On across multiple applications** - **Centralized user management and provisioning** - **Advanced session management and timeout policies** - **Detailed audit trails and compliance reporting** ### Multi-factor authentication (MFA) Identity providers that support MFA can automatically extend this security to your Customer Portal. When configured at the identity provider level, users will be required to complete additional authentication factors such as: - **SMS or email verification codes** - **Authenticator app tokens (Google Authenticator, Microsoft Authenticator, etc.)** - **Hardware security keys (FIDO2/WebAuthn)** - **Biometric authentication** - **Push notifications to mobile devices** The MFA enforcement is handled entirely by the identity provider, so no additional configuration is needed in Omnistrate once the identity provider is set up. ### Disabling default authentication For organizations that want to enforce SSO-only access, you can disable the default username/password authentication mechanism in your [Customer Portal settings](https://docs.omnistrate.com/tenant-management/customer-portal/index.md). This ensures that all users must authenticate through your configured identity providers, providing: - **Centralized access control** - **Consistent security policies** - **Simplified user management** - **Enhanced compliance with corporate security requirements** ## Troubleshooting ### Common Issues - **Invalid redirect URI**: Ensure the redirect URI in your identity provider matches exactly: `https://your-saas-domain.com/idp-auth` - **Scope errors**: Make sure you've included the required scopes: `openid email profile` - **Domain restrictions**: Check if your identity provider has domain restrictions that might block authentication - **HTTPS requirement**: Most identity providers require HTTPS for production use, Omnistrate ensures this is configured for the SaaS Portal, but if you are running the portal on your own environment (in particular in development environments) pls be mindful of this restriction. ### Testing your configuration Omnistrate performs best effort validation of your identity provider configuration while you enter the data. This includes checking the format of URLs, validating required fields, and testing connectivity to endpoints where possible. However, we strongly recommend thorough testing in a development environment before enabling identity providers in production. Important While Omnistrate validates configuration data during entry, the complete authentication flow can only be fully tested with real user interactions. Always perform comprehensive testing in development before enabling in production environments. For additional support, contact us at [support@omnistrate.com](mailto:support@omnistrate.com). # Tenant management overview Tenant management is the foundation of any successful SaaS platform, enabling you to organize, control, and scale your customers and operations effectively. This section covers the essential components for managing tenants, users, subscriptions, and the self-serve customer portal. ## What is a tenant? A **tenant** represents a distinct customer entity - typically an organization or company that subscribes to your services. Each tenant operates in isolation from others, ensuring data security and privacy. Tenants can range from individual users to large enterprises with multiple team members, each with their own access rights and resource allocations. In Omnistrate, tenants are managed through a combination of users, subscriptions, and organizational structures that provide both flexibility and security. ## Core Components ### Users and organizations [**User management**](https://docs.omnistrate.com/tenant-management/user-management/index.md) handles individual account holders who access your SaaS platform. Users can be external customers or internal team members, each with specific roles and permissions. Users are organized into organizations that represent the tenant structure, enabling team collaboration and access control. ### Subscriptions and plans [**Subscription management**](https://docs.omnistrate.com/tenant-management/subscription-management/index.md) governs how tenants access your services through formal agreements based on specific plans. Each subscription defines resource allocation, feature availability, service levels, and billing terms. Subscriptions can be auto-approved for streamlined onboarding or manually reviewed for controlled access. ### Automated self-service The [**customer portal**](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) provides a complete self-service experience that automates tenant onboarding and management. Customers can sign up, subscribe to plans, manage their resources, and handle billing - all without manual intervention. This automation reduces operational overhead while providing 24/7 availability for your customers. ## Advanced tenant management features ### Identity and access management - [**Identity providers**](https://docs.omnistrate.com/tenant-management/identity-providers/index.md) - Integrate with enterprise SSO solutions ### Network and security - [**Customer networks**](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md) - Dedicated network isolation per tenant - [**Private networking**](https://docs.omnistrate.com/tenant-management/private-networking/index.md) - Secure connectivity options ### Customization options - [**Build your own portal**](https://docs.omnistrate.com/tenant-management/build-your-own-portal/index.md) - Create custom tenant experiences using APIs # Private networking ## Introduction Private networking enables secure, direct connectivity between your customer's on-premises infrastructure or cloud networks and their deployed services in the cloud provider's account. This approach solves a critical challenge in cloud deployments: providing customers with secure access to their services without exposing them to the public internet. ### The Problem When deploying services in a cloud provider account, enterprise customers often need to: - Connect their existing infrastructure to deployed services securely - Avoid routing sensitive traffic over the public internet - Maintain network isolation and security boundaries - Comply with regulatory requirements that mandate private connectivity ### The Solution Private networking addresses these challenges by establishing secure, private connections between networks. This is achieved through various mechanisms: - [**VPC Peering**](#vpc-peering): Direct network connections between Virtual Private Clouds (VPCs) in the same or different cloud accounts - [**AWS PrivateLink**](#aws-privatelink): Provides private connectivity between VPCs and AWS services, enabling secure access without exposing traffic to the public internet. - [**Direct Connection**](#dedicated-connection): Dedicated, high-bandwidth private connections between customer networks and cloud provider infrastructure, bypassing the public internet for enhanced security and performance. These private networking options ensure that: - Traffic remains within private network boundaries - Network latency is minimized through direct routing - Security is enhanced by avoiding public internet exposure - Compliance requirements are met for data transmission ## VPC peering VPC peering allows you to connect two networks in different accounts or regions. This can be useful when you want to allow connectivity to resources from another account, without exposing them to public internet. It can be established from a network in the customer account to a network where customer resources are located. Because of that, this feature is limited only to BYOC, or Hosted SaaS with the "Customer Networks" feature enabled. In both those cases, each network is dedicated to a single customer. The VPC peering process can only be initiated by the customer, and the owner of the cloud network needs to accept the peering request. In the case of Hosted SaaS, that owner is the Service Provider. VPC peering for BYOC can be done solely by the customer. Warning VPC peering is currently not supported for Azure cloud provider. If you required this feature pls reach out to [support@omnistrate.com](mailto:support@omnistrate.com). Warning In order to make peering possible, CIDR blocks of source and target networks cannot overlap. This needs to be considered by the customer as they are the ones owning the source network and can also define the target network (which can either be imported existing network, created as a custom network with CIDR specified, or created as a default network by Omnistrate with default CIDR for a hosting cloud provider region). Note Use RFC1918 CIDR ranges on both sides of a private networking connection whenever possible (`10.0.0.0/8`, `172.16.0.0/12`, or `192.168.0.0/16`). Non-RFC1918 ranges are not recommended for Omnistrate-managed private connectivity and can complicate enterprise routing and peering. ### Step by Step instruction to configure VPC peering The following sections present the step-by-step guide to setting up VPC peering for AWS and GCP: #### AWS 1. Customer locates the network and collects details 1. On SaaS UI, navigate to "Tenant Management" -> "Customer Networks". 1. Locate the custom network you plan to peer. If it's not present, then such a network doesn't support VPC peering. 1. Note the following values: 1. **Native ID** - VPC ID of the network. 1. **CIDR** - CIDR range of the network (customer picked this value, but it can help with configuration). 1. **AWS Account ID** - AWS account ID that hosts the network (this will be one of your accounts). 1. **AWS Region** - AWS region of the network. 1. Share these values with your customer. 1. Customer uses the provided details to create a peering request. This can be done from within the AWS Console, CLI, or API. The next steps present an example of how to configure it using AWS Console: 1. In the AWS console navigate to: VPC -> Peering connections -> Create peering connection. 1. Select source VPC from the current account "VPC ID (Requester)". 1. Use **AWS Account ID** as a value for the "Another account" field. 1. Use **AWS Region** as a value for the "Another region" field (can be skipped if the region is matching). 1. Use **Native ID** as a value for the "VPC ID (Accepter)" field. 1. Configure the route table to route traffic to peered VPC. Configuration depends on **CIDR** of the target network (Accepter). For details, please see [AWS documentation](https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-routing.html). 1. Make sure "ACLs" on your VPC will not block outbound traffic. 1. Owner accepts a peering request (Service Provider or Customer, based on the Deployment Model) 1. In AWS console navigate to: VPC -> Peering connections. 1. Locate peering request. 1. Accept it. 1. Configure the VPC security group to allow traffic from peered VPC. For details, please see [AWS documentation](https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-security-groups.html). #### GCP 1. Customer locates the network and collects details 1. On SaaS UI, navigate to "Tenant Management" -> "Customer Networks". 1. Locate the custom network you plan to peer. If it's not present, then such a network doesn't support VPC peering. 1. Note the following values: 1. **Native ID** - GPC network identifier 1. **CIDR** - CIDR range of the network (customer picked this value, but it can help with configuration) 1. **GCP Project ID** - GCP project ID that hosts the network (this will be one of your projects) 1. Share these values with your customer 1. Customer uses the provided details to create a peering request. This can be done from within the GCP UI, CLI, or API. The next steps present an example of how to configure it using GCP UI: 1. On the GCP console navigate to: VPC Network -> VPC network details (of the network that will be peered) -> VPC network peering -> Add peering. 1. Pick a name for the peering connection that will help you identify it. 1. Select "In another project" and use **GCP Project ID** as a value for the "Project" field. 1. Use **Native ID** as a value for the "VPC network name" field. 1. Select "Import custom routes" and "Export custom routes". 1. Owner accepts a peering request (Service Provider or Customer, based on the Deployment Model) 1. On the GCP console navigate to: VPC Network -> VPC network details (of the **Native ID** network). 1. You should see the peering request. 1. Accept it. 1. Create firewall rules to allow traffic from the peered network. For details, please see [GCP documentation](https://cloud.google.com/vpc/docs/vpc-peering). ### Retrieve Network Information for VPC Peering from a Deployment Instance 1. Install the Omnistrate CTL from the [here](https://ctl.omnistrate.cloud/install/) 1. Identify the deployment cell associated with the deployment instance. ``` omnistrate-ctl instance describe -o json | jq '.DeploymentCellID' ``` 1. Obtain network details for the deployment cell identified in the previous step. ``` omnistrate-ctl custom-network list -f host_cluster_id: ``` ## AWS PrivateLink AWS PrivateLink provides secure, private connectivity between VPCs and AWS services or customer-hosted services without exposing traffic to the public internet. It creates a private endpoint within your VPC that acts as an entry point for traffic destined to supported AWS services or VPC endpoint services. ### Private Link Integration Using Omnistrate you can enable your customers to connect to their deployed services using AWS PrivateLink by automatically exposing VPC endpoint services that customers can reach from their own VPCs. This integration provides: - **Endpoint Service Creation**: Creates and manages VPC endpoint services for your service deployment - **Private Connectivity**: Customers can establish private connections to their services without exposing traffic to the public internet - **Cross-Account Access**: Secure connectivity between customer VPCs and Service Provider managed infrastructure - **Simplified Management**: Omnistrate handles the complexity of endpoint service configuration and management #### How It Works with Omnistrate 1. **Service Deployment**: When services are deployed through Omnistrate, VPC endpoint services are automatically created 1. **Endpoint Exposure**: Omnistrate exposes these endpoint services with appropriate permissions for customer access 1. **Customer Connection**: Customers create VPC endpoints in their own VPCs that connect to the Omnistrate-managed endpoint services 1. **Private Communication**: Traffic flows privately between customer VPCs and deployed services through the PrivateLink connection This approach allows customers to maintain private connectivity to their services while benefiting from Omnistrate's managed infrastructure and simplified networking configuration. For detailed implementation steps and support, reach out to [support@omnistrate.com](mailto:support@omnistrate.com). ## Dedicated Connection Dedicated Connection provides dedicated, private connections between customer on-premises infrastructure and cloud provider networks, bypassing the public internet entirely. This enterprise connectivity solution offers enhanced security, predictable performance, and reduced latency. ### How Dedicated Connection Works Direct link establishes a dedicated connection between customer premises and cloud provider infrastructure through: 1. **Physical Connection**: A dedicated circuit between customer network and cloud provider edge location 1. **Virtual Interfaces**: Multiple virtual connections (VLANs) over the physical connection 1. **BGP Routing**: Border Gateway Protocol exchanges routing information between networks 1. **Private Routing**: Traffic flows directly through the private connection ### Benefits - **Enhanced Security**: Traffic never traverses the public internet - **Predictable Performance**: Dedicated bandwidth with consistent latency - **Cost Optimization**: Reduced data transfer costs for high-volume workloads - **Compliance**: Meets regulatory requirements for private data transmission ### Setup Overview Setting up dedicated link requires coordination between customer, Service Provider, and cloud provider. The customer is responsible for network planning (determining bandwidth requirements and routing design), arranging physical connectivity to the cloud provider's edge location, configuring BGP routing protocols and policies, and ordering the dedicated circuit through the cloud provider. The Service Provider must accept the virtual interface in their account, configure network gateways and routing tables, set up security groups and access controls, and establish connection health monitoring. ### Cloud Provider Options - **AWS Direct Connect**: Dedicated connections to AWS infrastructure - **Azure ExpressRoute**: Private connections to Microsoft Azure - **Google Cloud Interconnect**: Dedicated connectivity to Google Cloud ### Key Considerations - **Redundancy**: Establish multiple connections for high availability - **Security**: Implement network segmentation and access controls - **Monitoring**: Track connection health and performance metrics - **Cost Management**: Right-size bandwidth and optimize data transfer costs For detailed implementation steps and support, reach out to [support@omnistrate.com](mailto:support@omnistrate.com). # Subscription management Omnistrate provides you with a powerful subscription management feature designed to simplify and automate the subscriptions for your SaaS Product. It allows you to configure subscription behaviors, such as enabling auto-approvals, or fine grained control such as manual approval, suspending/removing, or re-enabling subscriptions. You can gain a comprehensive view of all subscriptions, including detailed information on instances and user roles, ensuring complete visibility into your SaaS Products. ## What is a subscription? A subscription is a formal agreement between your customers and your SaaS platform that grants them access to your services based on a specific Plan. When customers subscribe to a Plan, they gain the ability to create and manage instances of your SaaS offering according to the resources, features, and limitations defined in that Plan. Each subscription represents: - **Access Rights**: Permissions to use your SaaS platform and its features - **Resource Allocation**: Compute, storage, and other resources as defined by the Plan - **Service Level**: The tier of service and support the customer receives - **Billing Agreement**: The pricing model and payment terms for using your platform ## How customers subscribe to a plan The subscription process is designed to be straightforward and secure for your customers: ### 1. Plan selection Customers browse the available Plans you've configured for your SaaS offering. Each Plan defines: - Resource limits (CPU, memory, storage) - Feature availability - Pricing structure - Support level ### 2. Subscription request When a customer selects a Plan, they initiate a subscription request through your customer portal or API. This request includes: - Selected Plan details - Customer account information - Billing preferences - Any custom configurations ### 3. Terms and conditions acceptance Before the subscription can be processed, customers must: - Review your platform's Terms of Service - Accept the Privacy Policy - Acknowledge the Service Level Agreement (SLA) - Agree to the billing terms and payment schedule This step ensures legal compliance and sets clear expectations for both parties. ### 4. Subscription approval Depending on your configuration, subscriptions can be: - **Auto-approved**: Instantly activated for immediate use - **Manually reviewed**: Queued for your team's approval before activation ### 5. Platform access Once approved and terms are accepted, customers can: - Access your SaaS platform - Create instances within their Plan limits - Manage their resources through the customer portal - Begin using your services immediately ### 6. Deployment instance management After subscription activation, customers can: - **Create Instances**: Launch new instances of your SaaS offering within their Plan limits - **Start/Stop Instances**: Temporarily halt running instances to save resources while preserving data - **Delete Instances**: Permanently remove instances they no longer need - **Monitor Performance**: Track instance performance, resource utilization, and health metrics - **Modify Configurations**: Adjust instance settings and parameters as allowed by their Plan - **Scale Resources**: Increase or decrease instance resources within Plan boundaries - **Access Management**: Control user access and permissions for their instances All these steps are automated and can be performed by customer in a self-served manner when enabling the [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) ## Subscription management operations Once customers subscribe to your SaaS Product, you will need to manage those subscriptions and perform operations as required on those subscriptions. Omnistrate enables the following operations: ### Subscription approval - auto/manual Whether you want tight control during the initial or beta launch of your SaaS, reviewing each subscription request before approving, or you want to streamline and automate the subscription approval process after launch, Omnistrate supports both approaches. You can configure it to meet your specific needs, providing you with the flexibility to manage subscriptions effectively at any stage. ### Review all subscriptions Review all of your subscriptions in detail with Omnistrate. You can see detailed list of all of your subscriptions with information such as status, users, roles and instances etc. This enables you to easily manage the subscriptions as well troubleshoot any issues that arise. Whether you need to check the status of a subscription, identify which users have access to specific subscriptions, or review the details of each instance, Omnistrate ensures you have all the information you need at your fingertips. ### Subscription suspension Subscriptions can be temporarily suspended in various situations, such as when billing issues arise or payments are pending. Suspending a subscription keeps it intact and customers can't create new instances, but their subscription details and any existing instances remain intact. This provides flexibility for customers who may have temporary budget constraints or seasonal needs, allowing them to easily resume their SaaS Product later without losing any progress. Subscription suspension affects access to the Plan for all customers on that subscription. User suspension is separate: it applies to an individual user account. A suspended user can still sign in and perform billing or payment-related activities, but cannot create, update, or manage deployment instances even if the subscription itself is active. For details, see [User Suspension](https://docs.omnistrate.com/tenant-management/user-management/#user-suspension). ### Reactivating a subscription Suspended subscriptions can be reactivated seamlessly, restoring full access to the SaaS Product and allowing customers to pick up right where they left off. This ensures a smooth and uninterrupted experience for customers who are ready to resume their subscription, eliminating any inconvenience. ### Terminating a subscription You can terminate subscriptions when needed. Whether ending a SaaS Product contract, responding to a customer cancellation request, enforcing policy compliance, or optimizing resource use, terminating subscriptions helps manage resources effectively. # User management Omnistrate provides you with comprehensive user management capabilities designed to simplify and automate user administration for your SaaS Product. You can gain complete visibility into all users across your platform, including detailed information on their subscriptions, roles, and activity, ensuring full control over your SaaS Product's user base. ## What is a user? A user is an individual account holder who has access to your SaaS platform through one or more subscriptions. Users can be either **external customers** who subscribe to and use your SaaS Product, or **internal customers** from your organization. Users represent the actual people who interact with your SaaS Products, create instances, and consume resources within the boundaries defined by their subscription plans. Each user represents: - **Email Address**: A unique email address that identifies the user - **Organization**: An organization that the users belong to - **Access Rights**: Permissions to access specific features and resources based on their subscriptions and user type - **Role Assignments**: Defined roles within each subscription that determine their level of access and capabilities ## How users sign up The user sign up process is built into the [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) and designed to be seamless and secure for your customers Note If you do not use the Customer Portal, you can register users using the API. ## Users and subscriptions Users access your SaaS Product by creating subscriptions to the plans you've defined. This relationship is fundamental to how your platform operates and how customers consume your services. To learn more about managing subscription see the [Subscription Management](https://docs.omnistrate.com/tenant-management/subscription-management/index.md) documentation. ## User management operations Once users have signed up and are actively using your SaaS Product, you will need to manage those users and perform various operations as required. Omnistrate enables the following user management operations: ### User suspension Users can be temporarily suspended in various situations, such as when policy violations occur, billing issues arise, or security concerns are identified. Suspending a user: - Preserves their ability to sign in - Preserves their account data and subscription information - Allows them to view billing information and perform payment-related activities, such as resolving outstanding invoices - Prevents them from creating new instances or performing DevOps operations on existing instances, including update, lifecycle management, upgrade, scale, delete, and other deployment-management actions - Maintains their subscription relationships for potential reactivation User suspension is designed to restrict operational access without blocking remediation. For example, a suspended user can still sign in to resolve billing issues, but cannot create or manage deployments until the suspension is removed. Suspension is different from deleting a user or disabling an authentication method. A suspended user account still exists and can authenticate; the suspension is enforced when the user attempts restricted operations. ### User reactivation Suspended users can be reactivated seamlessly, restoring full access to their subscriptions and services. Reactivation: - Restores complete platform access - Allows users to resume their previous activities - Maintains continuity of their subscription and instance data - Ensures a smooth experience when issues are resolved This ensures users can quickly return to productivity once any concerns have been addressed. ### User verification User verification allows you to confirm the authenticity and legitimacy of user accounts. This is useful in the case where you look to validate users email with a different mechanism that the out of the box emails sent by the Customer Portal. ### User export The user export functionality allows you to extract comprehensive user data for reporting purposes. ### User deletion Users can be permanently deleted when necessary. This operation is typically used for: - Fulfilling user requests for account deletion (right to be forgotten) - Removing inactive or abandoned accounts - Cleaning up test or duplicate accounts **Important**: User deletion is permanent and irreversible. All associated data will be permanently removed from the system. # DevOps Guides # Environments ## Environments Overview Environments in Omnistrate provide isolated, configurable deployment contexts that enable you to manage different stages of your software development lifecycle. They serve as the foundation for implementing robust deployment pipelines, testing strategies, and production operations. Each environment serves a specific purpose in the software development lifecycle, allowing for controlled testing and validation of code changes before they are promoted to the next stage or, ultimately, to production. Omnistrate automates the whole process to deliver an efficient and reliable software delivery pipeline from promotions across environments, access control etc. Omnistrate environments are logical groupings of infrastructure and configuration that allow you to: - **Isolate Workloads**: Separate development, testing, staging, and production deployments - **Control Access**: Implement role-based access controls per environment - **Enable Promotions**: Facilitate automated and manual deployment promotions between environments - **Manage Configurations**: Apply environment-specific configuration - **Ensure Consistency**: Maintain configuration parity while allowing environment-specific customizations ## Environment Types ### Development Environment This is where software developers work on new features or bug fixes. It typically mirrors the developer's local development environment and may include development databases, code repositories, and development tools. ### Testing Environment The testing environment is used for quality assurance and automated testing. It is designed to mimic the production environment as closely as possible to identify potential issues before code is deployed to production. ### Production Environment The production environment is where the live application or service is hosted and accessed by end-users. It is the final destination for code changes after they have been tested and approved in previous environments. ### Sandbox Environment Developers may create temporary environments for working on specific features or bug fixes in isolation before merging their changes into the main development branch. ## Access Control per Environment You can choose to enable specific [access controls](https://docs.omnistrate.com/governance-guides/omnistrate-rbac/index.md) to restrict who can enable promotions across environments. See more information about [promotion pipelines](https://docs.omnistrate.com/dev-ops-guides/pipelines/index.md) ## Going live with your SaaS Product Production environments can be made Public to make your SaaS Product available to your customers. Omnistrate provides [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) - ready to use Web Application - for your end customers. Once you have made an environment public and configured Customer Portal, your SaaS Product will be available on your custom domain for your customers to use. To configure custom domain, please refer to the [Customer Portal guidelines](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) # DevOps Guide Overview Omnistrate's DevOps capabilities provide comprehensive tools and workflows to streamline your software development lifecycle, from initial development through production deployment. These guide covers the essential components for building robust, automated, and secure deployment processes. ## Core DevOps Components ### [Environments](https://docs.omnistrate.com/dev-ops-guides/environments/index.md) Environments form the foundation of your deployment strategy, providing isolated, configurable contexts for different stages of your software development lifecycle. ### [Pipelines](https://docs.omnistrate.com/dev-ops-guides/pipelines/index.md) Pipelines orchestrate the automated promotion of changes between environments, ensuring reliable and consistent deployments through validation gates and approval workflows. ### [Upgrades](https://docs.omnistrate.com/dev-ops-guides/upgrades/index.md) Systematic approaches to upgrading applications, infrastructure, and dependencies across environments. ### [Secrets](https://docs.omnistrate.com/dev-ops-guides/secrets/index.md) Secure management of sensitive information across environments. # Pipelines ## Pipelines overview Omnistrate's pipeline enables you to implement robust continuous deployment (CD) processes that streamlines the promotion of changes from development through to production. At the heart of these pipelines is the concept of **promotions** - the controlled movement of code, configurations, and deployments between [environments](https://docs.omnistrate.com/dev-ops-guides/environments/index.md) to ensure reliable and efficient software delivery for your SaaS product. Omnistrate pipelines are built around promotion-driven workflows that leverage [environments](https://docs.omnistrate.com/dev-ops-guides/environments/index.md) to create automated deployment processes: - **Automate Promotions**: Seamlessly move configurations, code, and deployments between environments with built-in validation - **Enforce Validation**: Implement approval gates, automated testing, and quality checks at each promotion stage - **Maintain Consistency**: Ensure configuration parity and deployment consistency across environment stages - **Enable Safe Rollbacks**: Provide quick recovery mechanisms when promotions encounter issues - **Track Promotion History**: Maintain comprehensive audit trails of all promotion activities ## Promotions Promotions are the core mechanism that drives Omnistrate pipelines, enabling the controlled and validated movement of changes through your deployment environments. ### What Are Promotions? A promotion in Omnistrate represents the process of moving validated changes from one environment to the next in your deployment pipeline. Promotions can include: - **Code Deployments**: Moving application code and container images between environments - **Configuration Changes**: Promoting environment-specific configurations and parameters - **Infrastructure Updates**: Applying infrastructure changes and resource modifications Note Omnistrate treats infrastructure, software and configuration as a single configuration unit, allowing you to promote all changes seamlessly across environments. ## Pipeline Architecture Omnistrate supports flexible pipeline architectures based on your deployment needs: ### Standard Pipeline Flow ### Advanced Pipeline Flow ## Role Based Access (RBAC) Omnistrate RBAC to enforce access control restricting who can perform a promotion, see [here](https://docs.omnistrate.com/governance-guides/omnistrate-rbac/index.md) ## Dry Run Releases A dry run release allows you to validate Plan changes before deploying them to an environment. When you perform a dry run, Omnistrate simulates the release process — checking configuration validity, resource compatibility, and parameter consistency — without applying any changes to running deployment instances. ### Dry run with CTL You can perform a dry run build using the Omnistrate CTL with the `-d` or `--dry-run` flag: ``` omnistrate-ctl build --file omnistrate-compose.yaml --product-name "My SaaS Product" --dry-run ``` This simulates building the SaaS Product without actually creating resources, allowing you to validate your specification before applying changes. Note Dry run releases do not create new versions or modify any existing deployment instances. They are purely a validation mechanism. ## Continuous Integration Omnistrate supports continuous integration workflows that enable automated building, testing, and deployment of your applications. The CI process typically involves building Docker images, running tests, and publishing versioned images to Omnistrate using GitHub Actions or other CI/CD platforms. For a complete guide on setting up a continuous integration process with Omnistrate, including GitHub Actions configuration, automated testing, and best practices for image versioning, see the [CI/CD example repository](https://github.com/omnistrate-community/ci-cd-example). To integrate with GitHub Actions, you can use the [Setup Omnistrate CTL action](https://github.com/marketplace/actions/setup-omnistrate-ctl) from the GitHub Marketplace to easily configure the Omnistrate CTL in your workflows. Tip Use immutable image references such as versioned tags or digests rather than fixed tags like `latest` to ensure predictable deployments and maintain fine-grained control over resource instance upgrades. Tip Avoid mutable tags such as `latest` for production rollouts, especially for sidecars and helper containers. Mutable tags make promotions harder to reason about and can lead to stale cached images being reused unexpectedly across upgrades. # Secrets ## Overview Omnistrate allows you to include sensitive information in your deployment specifications through the use of secrets. A secret is defined as a name-value pair. Secret names can be used as placeholders in service definitions that are replaced with their respective values during a deployment. You can use secrets in various places in your Helm Charts, Kubernetes Operators, Kustomize templates, Terraform templates, OpenTofu templates, and Docker Compose configurations. Secrets are defined at the **environment type level** (Dev, Stage, Prod, etc.). This means you can use different values for the same secret depending on what type of environment the deployment is launched in. For example, you could define a `dbPassword` secret that has the value of your development database password in your Dev environments and the value of your production database password in Prod environments. Info All secret values are stored as string type. ## Creating Secrets You can create and manage secrets using the Omnistrate UI, CTL, or API. ### Using the UI Navigate to the **Secrets** section in the Omnistrate console. You can create, update, and delete secrets from the environment settings. ### Using the CTL Use the following CTL commands to manage secrets: ``` # Create a secret omnistrate-ctl secret create --name dbPassword --value "my-secure-password" --environment-type PROD # List secrets omnistrate-ctl secret list --environment-type PROD # Update a secret omnistrate-ctl secret update --name dbPassword --value "new-secure-password" --environment-type PROD # Delete a secret omnistrate-ctl secret delete --name dbPassword --environment-type PROD ``` ### Using the API You can also manage secrets programmatically through the Omnistrate API. Refer to the [API documentation](https://docs.omnistrate.com/api/api-resources/index.md) for the full list of secret management endpoints. ## Secret Syntax Secrets use two syntax formats depending on the context: | Syntax | Context | Example | | ----------------------------- | ----------------------------------------------------- | ------------------------- | | `$secret.` | Docker Compose environment variables, Operator config | `$secret.dbPassword` | | `$secret.` in templates | Helm chart layered values templates | `$secret.GH_ACCESS_TOKEN` | ## Usage in Docker Compose Use the `$secret.` syntax in environment variables or configuration values: ``` services: postgres: image: postgres:16 environment: - POSTGRESQL_DATABASE=$var.dbDatabase - POSTGRESQL_USERNAME=$var.dbUsername - POSTGRESQL_PASSWORD=$secret.dbPassword ``` ## Usage in Helm Charts ### Chart values Use the `$secret.` syntax when referencing secrets in `chartValues`: ``` services: - name: Example Service helmChartConfiguration: chartName: example-chart chartVersion: 0.1.1 chartRepoName: example-chart-repo chartRepoURL: https://raw.githubusercontent.com/omnistrate-community/example-chart-repo/main chartValues: auth: database: $var.dbDatabase username: $var.dbUsername password: $secret.dbPassword ``` ### Layered chart values When using [layered chart values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md), secrets are available as template variables. This is useful for providing private repository access tokens and other credentials: ``` chartValueLayers: - scope: cloudProviderRegion: "*" valuesFile: gitConfiguration: repositoryUrl: "https://github.com/your-org/helm-configs.git" reference: "refs/heads/main" accessToken: "$secret.GH_ACCESS_TOKEN" path: "values.yaml" ``` For more details on template variables in layered chart values, see the [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) guide. ## Usage in Operators Secrets can be referenced in Operator configurations using the `$secret.` syntax, the same way they are used in Docker Compose: ``` services: - name: My Operator operatorCRDConfiguration: crdName: my-operator helmChartConfiguration: chartName: my-operator-chart chartValues: operatorSecret: $secret.operatorApiKey ``` ## Usage in Kustomize and Terraform Secrets are available in Kustomize and Terraform/OpenTofu templates using the `$secret.` syntax. Use them anywhere you need to inject sensitive values into your deployment specifications. ## Kubernetes Secrets as Cell Amenities You can deploy Kubernetes `Secret` resources directly to deployment cells using [cell amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md). This is useful for pre-provisioning secrets that your workloads need at the cluster level: ``` customAmenities: - name: app-secrets description: Application secrets type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: Secret metadata: name: app-secrets namespace: default type: Opaque stringData: API_KEY: my-api-key ``` For more details, see the [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) guide. ## Best Practices - **Use environment-scoped secrets** to maintain separate credentials for Dev, Stage, and Prod environments. - **Use the `password` API param type** for any user-facing sensitive inputs to ensure values are masked. - **Avoid hardcoding secrets** in compose files or Helm values — always use `$secret.` placeholders. - **Rotate secrets regularly** by updating secret values through the UI, CTL, or API. # Managing Upgrades ## Managing Upgrade with Omnistrate Managing upgrades refers to the process of applying updates, fixes, or patches to the software and configurations including infrastructure upgrades. More specifically: - Upgrading software images and CVEs - Upgrading infrastructure - Changing configuration Info Like other features, upgrading is NOT a simple lipstick on stateful sets but something carefully designed for stateful systems. We overcome the common challenges to have per node configuration, provide visibility into the upgrade process, provide customization etc. Moreover, upgrading 1 machine is different from upgrading the whole fleet. We have built an end-to-end system that allows you to manage different types of upgrades from images, infrastructure, operating system to configuration across all of your customers and have the control/flexibility to meet the needs of different customers at the same time without manual toil. ## Types of Upgrades Let's look into each of scenarios and understand some of the use cases better: ### Upgrading software There might be many reasons to patch software: - Security Updates: to address security vulnerabilities and apply security patches promptly - Bug Fixes: to address these issues, improving the overall reliability of the service - Feature Enhancements: to introduce new features or capabilities, enhancing its functionality and usability - Performance Enhancements: to better response times and infrastructure utilization ### Upgrading infrastructure You may have to upgrade infrastructure due to several reasons: - Replace end of life hardware - Improve price to performance ratio - Feature enhancements ### Changing configuration You may have to change configuration: - Better defaults - Add or remove configuration due to feature enhancements/depreciation - Evolve your SaaS experience ## Upgrade Lifecycle Now, let's look into how Omnistrate helps address the Upgrading across different scenarios. The upgrade lifecycle is divided into 3 phases: ### Make changes At Omnistrate, all the changes across software, configuration or infrastructure are versioned. As soon as you release a new version, that version is locked and immutable. To better manage different versions, each version has an associated state. Today, there are 3 possible states: - Active: Active version in the system that are valid versions in the system. They can be marked as Preferred or can be an upgrade target for the existing customer resource instances. - Preferred: Preferred version in the system that will be used by existing and newer customers for any new usage. There can only be 1 and only 1 Preferred version in the system at any given time. If you promote another Active version as Preferred, the previous Preferred version state will automatically change back to Active. - Deprecated: Deprecated versions in the system that are no longer valid. By definition, there can't be any customer resource instances associated with the deprecated versions. Every time, you make a change to your SaaS, Omnistrate will automatically create a new unreleased version or append changes to a unreleased version. Once the changes are released, that version is marked as released and can't be changed anymore. By default, the newer versions are released as Active that can be marked as Preferred at any time. You may want to run some tests before marking it as Preferred to be a default version used by your customers. You can see the list of all the plan versions on the `View Plan Versions` screen. #### How to make the changes? If you have already created a SaaS, you can use CTL/API/UI to make further changes. If you want to update the service with a new version of docker compose, CTL is the recommended way. You can also visit "Build Service" in UI/API. Your SaaS Product may have several Plans for your customers. You have to then select the Plan that you want to update by clicking "Build Plan", see below: You will then see a DAG of different Resources. You can add/remove a Resource or update the dependencies or modify different elements of the Resource. To learn more about different elements of the Resource, see [here](https://docs.omnistrate.com/build-guides/resource/#elements-of-resource) Note Note that Omnistrate versions all of your control plane (or your SaaS) changes so that you can keep track, audit, rollback your changes. In addition, you can define the migration strategy for the existing resources owned by different tenants. Here is some more information if you are trying to modify different elements of the Resource (also see highlighted in green in the picture below): - Updating API parameters: To make any changes to API parameters, select the component and visit API parameters section to view/modify/delete the corresponding API parameters. - Updating configuration: To make any changes to your Resource configuration, select the component and visit the environment variables section to view/modify/delete them. We support two kinds of configuration: 1/ static 2/ dynamic. For Dynamic configuration, you can wire any API parameter as the value to your configuration. This will allow you to customize the stack for your customers and serve their needs respectively. - Updating images: To make any changes to image, visit the infrastructure tab and make the required changes - Updating infrastructure: To make any changes to infrastructure, visit the Compute, Network and Storage tabs and make the required changes. Note that you can also configure your infrastructure dynamically in the same way as service configuration (discussed above) by wiring API parameter to your infrastructure configuration - Updating action hooks: To make any changes to action hooks, select the component and visit the action hooks section. Action hooks are at two levels: 1/ nodes 2/ cluster. Please select the scope, then select the type of action hook that you want to update. - Configuring capabilities: To add/remove capabilities, click on the Capabilities tab and make the required changes - Adding/Removing integrations: To add/remove integrations, click on the Integrations tab and make the required changes - Updating service dependencies: You can define service dependencies from the dependencies section. Note that, if Resource B depends on A, then you have to map all required API params of B from API params of A. Once the changes are made, you have to release them for them to take effect. Also - if you want your new customers to see the latest changes, you have to make the release or the new version as Preferred. ## Schedule upgrades Over time, there might be several Preferred versions and you may want to upgrade customer resource instances running on older version to use newer version. Omnistrate lets you define upgrade paths as one-off patch or upgrade all resource instances on a given version to a different valid version. Once upgraded, the older version can be deprecated over time to manage the number of active versions in the fleet. ### Manage scheduled and ongoing upgrades The upgrade process can be requested to start immediately or scheduled for a specific time in the future. Note In case of scheduled upgrades, SaaS Providers and end customers will receive notifications about incoming upgrades, at the time of scheduling and 1 hour before the upgrade process starts. While dependent on the current lifecycle of the process, you can manage ongoing upgrades: Omnistrate provides robust tools to oversee and control scheduled upgrades, ensuring seamless tracking and execution. You can: - **Monitor upgrade progress**: You can look at the details to get a specific status of any resource instance that's part of this upgrade. - **Maintain upgrade lifecycle**: Cancel scheduled or cancel/resume upgrades depending on the current path lifecycle. - **Handle failures in a batch**: We roll patches in batches to reduce blast radius and run instance upgrades sequentially. If one or multiple instances fail during a batch upgrade, the process automatically pauses and triggers a notification Alert. This allows your service operator to assess the situation and decide whether to proceed with or cancel the remaining batches. Omnistrate automatically figures out the deployment plan and roll it out to your fleet. In addition, Omnistrate act as a central nervous system for all the activities in your fleet and coordinates them appropriately to avoid interference across multiple operations on the same resources. Separately, Omnistrate follows the dependency structure between your Resources to upgrade them in the right order. As an example, if A depends on B, then we will first upgrade dependencies and then the parent resource to ensure minimal disruption and maintain high availability throughout the upgrade process. If you notice any issues during the upgrade, you may pause or cancel a scheduled and pending upgrades. Note If you pause or cancel an in-progress upgrade path, the change will only affect resource instances that are in a pending or scheduled state. Any instances already in progress will continue without interruption. If you really want to cancel upgrades to resource instances in progress, you will have to pause or cancel the corresponding workflows ## Retrieve Plan Specification You can retrieve the specification for any Plan version through the UI or API. This is useful for auditing changes between versions, debugging configuration issues, or restoring a previous version's configuration. Note This feature is only available for Plans created with a Plan specification or compose spec. Plans created through the UI wizard do not have a retrievable specification. To retrieve the specification, navigate to your SaaS Product, select the Plan, and open the **Version History** tab. Select a version to view its full specification. To retrieve the specification programmatically, use the `GET /2022-09-01-00/service/{serviceId}/service-plan/{planId}/compose-spec` endpoint. Refer to the [API documentation](https://api.omnistrate.cloud/docs/external/) for full details. # Infrastructure Guides # Blob storage ## How to mount blob storage The blob storage feature allows you to define blob storage(s) and their mount paths across Resources. Omnistrate will manage the lifecycle of such storage volumes (e.g: creation and deletion together with instance) as well as mounting them as a local volume according to the configuration. Using a blob storage provides virtually unlimited space for your application where you pay only for the storage you use. Note This feature is not available for the Hosted SaaS model. ## Defining blob storage You can define blob storage(s) in the compose specification as follows: ``` services: serviceA: volumes: - source: bucket_data target: /mnt/blob-metadata type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::S3 gcp: clusterStorageType: GCP::GCS azure: clusterStorageType: AZURE::BLOB_STORAGE volumes: bucket_data: driver: blob ``` - `bucket_data` is the name of the volume, you can opt for a name of your choice. - `driver: blob` is the driver type required to enable using Blob storage. - `x-omnistrate-storage` defines the type of storage for each cloud provider. This configuration will provision a new Blob storage (bucket/container) for each instance deployed. On the pod itself, it will appear just like any other volume. The underlying storage driver does not restrict or synchronize access, so your application itself needs to make sure it doesn't end up overriding data. ## Mounting blob storage in multiple resources The blob can be mounted in different pods: ``` services: serviceA: volumes: - source: bucket_data target: /mnt/blob-metadata type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::S3 gcp: clusterStorageType: GCP::GCS azure: clusterStorageType: AZURE::BLOB_STORAGE serviceB: volumes: - source: bucket_data target: /mnt/blob-metadata type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::S3 gcp: clusterStorageType: GCP::GCS azure: clusterStorageType: AZURE::BLOB_STORAGE volumes: bucket_data: driver: blob ``` > **Note:** Blob storage is currently only available for GCP GCS and AWS S3. Contact [support@omnistrate.com](mailto:support@omnistrate.com) for other options. # Compute Management Omnistrate provides comprehensive compute management capabilities that abstract away cloud complexities while maintaining full control over your infrastructure. The platform manages compute resources intelligently based on your tenancy model, deployment requirements, and multi-cloud strategy. ## How Omnistrate Manages Compute Omnistrate's compute management is built around several key principles: - **Tenant-Aware**: Compute resources are allocated and managed per-tenant based on your chosen tenancy model - **ACID-Compliant**: All compute operations are fully ACID-compliant, ensuring no leaked resources - **Versioned**: All compute configurations and changes are versioned for rollback and audit capabilities - **Multi-Cloud**: Abstracts cloud provider differences while preserving cloud-specific optimizations - **Day-2 Operations**: Automates ongoing compute operations, not just initial provisioning ## Compute Management by Tenancy Type ### Dedicated Tenancy Compute Management In **Dedicated Tenancy** (`OMNISTRATE_DEDICATED_TENANCY`), each tenant receives completely isolated compute infrastructure with dedicated virtual machines, storage, and network resources. #### How Dedicated Compute Works **Infrastructure Isolation**: Each tenant deployment gets its own dedicated virtual machines with no sharing at the infrastructure level. Omnistrate provisions, manages, and scales these resources independently per tenant. **Resource Allocation**: Compute resources are allocated based on: - Service requirements defined in your compose specification - Customer-specific resource parameters - Performance and compliance requirements - Geographic placement preferences **Scaling Behavior**: When a dedicated tenant needs more compute capacity, Omnistrate: 1. Provisions additional virtual machines in the same deployment cell 1. Configures networking and storage connectivity 1. Deploys and configures your service components 1. Updates load balancing and service discovery **Configuration Example**: ``` x-omnistrate-service-plan: tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' services: app: x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: m5.large - cloudProvider: gcp name: n1-standard-2 - cloudProvider: azure name: Standard_D2s_v3 ``` **Compute Features for Dedicated Tenancy**: - Custom instance types per cloud provider - Dedicated CPU and memory allocation - Isolated storage volumes - GPU support for AI/ML workloads ### Multi-Tenancy Compute Management In **Multi-Tenancy** (`OMNISTRATE_MULTI_TENANCY`), multiple tenant instances share underlying compute infrastructure through intelligent bin-packing and resource optimization. #### How Multi-Tenant Compute Works **Intelligent Bin-Packing**: Omnistrate automatically places multiple tenant workloads on shared virtual machines based on: - CPU and memory requirements - Resource utilization patterns - Performance isolation requirements **Resource Isolation**: While infrastructure is shared, each tenant gets: - Dedicated CPU and memory limits/requests - Isolated container runtime environments - Network-level isolation through Kubernetes namespaces - Storage isolation through persistent volume claims **Dynamic Scaling**: Omnistrate continuously monitors resource utilization and: - Adds new virtual machines when capacity is exceeded - Scales down unused capacity to optimize costs - Maintains performance isolation during scaling events **Configuration Example**: ``` x-omnistrate-service-plan: tenancyType: 'OMNISTRATE_MULTI_TENANCY' services: app: platform: "linux/arm64" # Use ARM64 for cost optimization deploy: resources: limits: cpus: '2.0' memory: 4G requests: cpus: '0.5' memory: 1G ``` **Compute Features for Multi-Tenancy**: - Automatic resource optimization and bin-packing - Configurable CPU and memory limits per tenant - Shared infrastructure with logical isolation - Cost-effective resource utilization - Automatic scaling based on aggregate demand - Platform architecture support (x86_64, ARM64) ## Multi-Cloud Compute Options Omnistrate provides comprehensive multi-cloud compute capabilities, allowing you to deploy and manage compute resources across multiple cloud providers seamlessly. ### Supported Cloud Providers **Amazon Web Services (AWS)** - Full range of EC2 instance types - Graviton (ARM64) and Intel/AMD (x86_64) processors - GPU instances for AI/ML workloads (P3, P4, G4, etc.) - Spot instances for cost optimization - Regional deployment across all AWS regions **Google Cloud Platform (GCP)** - Compute Engine instance families - Custom machine types for precise resource allocation - GPU accelerators (NVIDIA T4, V100, A100) - Preemptible instances for cost savings - Regional deployment across GCP regions **Microsoft Azure** - Virtual Machine families - ARM64 and Intel/AMD processors - Azure Spot Virtual Machines - GPU-enabled instances for compute-intensive workloads - Regional deployment across Azure regions ### Instance Type Selection Specify compute requirements per cloud provider: ``` services: database: x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: r5.xlarge # Memory-optimized for database - cloudProvider: gcp name: n1-highmem-4 # High memory equivalent - cloudProvider: azure name: Standard_E4s_v3 # Memory-optimized equivalent ``` ### Per-cloud instance type overrides When no explicit instance type is configured, Omnistrate uses your Resource's `platform` value and resource requests to select suitable compute for each cloud provider. For example, `platform: linux/arm64` directs Omnistrate to select ARM64-compatible defaults for supported clouds. You can override the selected instance type for each cloud provider with `x-omnistrate-compute.instanceTypes`: ``` services: app: image: ghcr.io/example/app:latest platform: linux/arm64 x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: r7g.large - cloudProvider: gcp name: t2a-standard-4 - cloudProvider: azure name: Standard_E2ps_v6 deploy: resources: requests: cpus: '0.5' memory: 1G ``` Explicit instance type overrides take precedence over platform-based defaults for the cloud provider where they are defined. This also allows mixed per-cloud choices, such as AWS on ARM and GCP on x86: ``` services: app: image: ghcr.io/example/app:latest platform: linux/arm64 x-omnistrate-compute: instanceTypes: - cloudProvider: aws name: r7g.large # ARM64 - cloudProvider: gcp name: n2-highmem-2 # x86_64 ``` If you select an instance type that has node-local NVMe, see [Local NVMe Storage](https://docs.omnistrate.com/infra-guides/local-nvme-storage/index.md) — on AWS and Azure those disks are prepared and mounted automatically, and on GCP local SSD is configured through `configurationOverrides`. Instance type overrides are resource-local. Omnistrate does not inherit `instanceTypes` from a parent resource. In a large resource tree, add `instanceTypes` only to compute-backed resources that need a specific override. If several compute-backed resources need the same selected values, define API parameters once on the parent resource and pass them to child resources with `parameterDependencyMap`. Each child resource still needs its own `x-omnistrate-compute.instanceTypes` entries that reference those API parameters. ``` services: app: image: omnistrate/noop x-omnistrate-api-params: - key: awsInstanceType description: AWS instance type name: AWS Instance Type type: String modifiable: false required: false export: true defaultValue: r7g.large parameterDependencyMap: api: awsInstanceType worker: awsInstanceType - key: gcpInstanceType description: GCP instance type name: GCP Instance Type type: String modifiable: false required: false export: true defaultValue: n2-highmem-2 parameterDependencyMap: api: gcpInstanceType worker: gcpInstanceType depends_on: - api - worker api: image: ghcr.io/example/api:latest x-omnistrate-api-params: - key: awsInstanceType description: AWS instance type name: AWS Instance Type type: String modifiable: false required: true export: false - key: gcpInstanceType description: GCP instance type name: GCP Instance Type type: String modifiable: false required: true export: false x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: awsInstanceType - cloudProvider: gcp apiParam: gcpInstanceType worker: image: ghcr.io/example/worker:latest x-omnistrate-api-params: - key: awsInstanceType description: AWS instance type name: AWS Instance Type type: String modifiable: false required: true export: false - key: gcpInstanceType description: GCP instance type name: GCP Instance Type type: String modifiable: false required: true export: false x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: awsInstanceType - cloudProvider: gcp apiParam: gcpInstanceType ``` `parameterDependencyMap` applies only to the direct child resources listed in the map. If your dependency graph has multiple levels, map the parameter to each resource that needs the value. When mixing architectures across clouds, publish a multi-architecture container image or ensure each selected instance type can run the container image you provide. An ARM64-only container image will not run on an x86_64 node pool, even if `platform: linux/arm64` is set on the Resource. # Custom Deployment Cell Placement ## Deployment Cells Overview Custom Deployment Cell Placement allows you to control how many deployments can be co-located on the same host cluster (deployment cell / Kubernetes cluster). This provides fine-grained control over deployment isolation and resource allocation for your service instances. The `CUSTOM_DEPLOYMENT_CELL_PLACEMENT` feature enables you to specify the maximum number of deployment instances that can be placed on a single deployment cell. This is particularly useful for: - **Deployment isolation**: Setting `maximumDeploymentsPerCell: 1` ensures dedicated host clusters for each deployment - **Resource optimization**: Higher values allow resource sharing across multiple deployments - **Compliance requirements**: Meeting security standards that require isolation - **Performance guarantees**: Ensuring predictable performance by limiting resource contention Warning Setting `maximumDeploymentsPerCell: 1` can significantly increase infrastructure costs as it provisions dedicated host clusters for each deployment. Empty host clusters are automatically deleted after 24 hours of inactivity. Note With `maximumDeploymentsPerCell: 1`, the first deployment time will include host cluster creation time (typically 5-10 minutes). While empty host clusters with deleted deployments can be reused if available. ## Configuration of Custom Deployment Cells Custom Deployment Cell Placement can be configured in both compose spec and Plan spec files: ### Compose Spec Configuration Add the feature under `x-omnistrate-service-plan`: ``` version: '3.9' x-omnistrate-service-plan: name: "My Plan" features: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 1 ``` ### Plan Spec Configuration Add the feature directly under `features`: ``` name: "My Plan" features: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 1 ``` ### Configuration Options | Parameter | Type | Required | Description | | --------------------------- | ------- | -------- | ------------------------------------------------------------------------------- | | `maximumDeploymentsPerCell` | integer | Yes | Maximum number of deployments allowed per host cluster. Must be greater than 0. | ### Valid Values - `1`: Dedicated host cluster per deployment (complete isolation) - `2-N`: Shared host clusters with specified maximum deployments per cluster ## Examples of Custom Deployment Cell Configuration ### Compose Spec Example: Dedicated Host Clusters ``` version: '3.9' x-omnistrate-service-plan: name: "MySQL Service" features: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 1 services: mysql: image: mysql:8.0 ports: - "3306" environment: - MYSQL_ROOT_PASSWORD=secret ``` ### Plan Spec Example: Shared Host Clusters ``` name: "PostgreSQL Service" features: CUSTOM_DEPLOYMENT_CELL_PLACEMENT: maximumDeploymentsPerCell: 3 services: - name: postgres helmChartConfiguration: chartName: postgresql chartVersion: "12.1.7" chartRepository: https://charts.bitnami.com/bitnami ``` ## How Custom Deployment Cells works When placing a new deployment, the system: 1. Finds existing host clusters with available capacity (picks the one with least room left) 1. If no suitable cluster exists, provisions a new one 1. Considers region, cloud provider, and resource requirements 1. Respects the `maximumDeploymentsPerCell` constraint Host clusters are automatically created when needed and can be reused after deployments are removed. Empty host clusters are automatically deleted after 24 hours of inactivity. When multiple services share the same cloud account, they can share host clusters while respecting each service's individual constraints. For example, if Service A has `maximumDeploymentsPerCell: 2` and Service B has `maximumDeploymentsPerCell: 3`, a shared host cluster will allow up to 2 deployments from Service A and up to 3 deployments from Service B simultaneously. ## Limitations of Custom Deployment Cells - Not compatible with multi-tenant offerings - Not compatible with Serverless capability - Can only be modified or enabled/disabled when no deployments exist for the service - `maximumDeploymentsPerCell` must be greater than 0 # Endpoint aliases ## Configuring Endpoint aliases Omnistrate additionally enables users to assign specific aliases to any deployment or resource instance endpoints. By configuring DNS records, users can map their deployment endpoints to these aliases. These are configured per Resource. Eg: If you have two Resources DB and Redis in your SaaS architecture, you can allow your users to configure a custom DNS for each of these Resources separately. Info Please note that endpoint aliases feature is only available in the enterprise plan for now. The feature offers the capability to configure the deployments with specific domain, providing enhanced branding, customization, and control over their infrastructure. The endpoint aliases feature supports SSL/TLS encryption for secure communication between clients and the configured endpoint. ### Compose Spec Configuration Users can enable the endpoint aliases feature in the compose spec: ``` services: web: image: nginx x-omnistrate-capabilities: customDNS: targetPort: 80 ``` ### Plan Spec Configuration If you're using the Plan spec for deploying a Helm chart, Kustomize resource or an Operator CRD, you can enable it in the Plan spec: ``` services: - name: web capabilities: customDNS: TargetKubernetesService: TargetName: my-service TargetPort: 80 ``` `targetPort` is the port number where your http service is listening on. `TargetKubernetesService` is the target Kubernetes service to which the alias should be mapped to. ### Enabling custom DNS with HTTPS load balancer You can also enable custom DNS at the HTTPS (L7) load balancer level by setting `enableCustomDNS: true` on the load balancer configuration. This applies the custom DNS alias to the load balancer endpoint rather than to an individual Resource endpoint. Warning Enabling custom DNS on the HTTPS load balancer and enabling the `customDNS` capability on a Resource are mutually exclusive. Only one of these options can be active at a time. If you enable custom DNS on the load balancer, do not also configure the `customDNS` capability on the associated Resource, and vice versa. #### Compose Spec Configuration ``` x-omnistrate-load-balancer: https: - name: api-gateway enableCustomDNS: true paths: - associatedResourceKey: gateway path: / ``` #### Plan Spec Configuration ``` loadBalancers: https: - name: api-gateway enableCustomDNS: true paths: - associatedResourceKey: gateway path: / backendPort: 80 ``` `enableCustomDNS` enables custom DNS on the L7 load balancer, allowing your customers to configure a domain alias for the load balancer endpoint. ## Custom DNS as an Input Parameter When you enable the `customDNS` capability on a Resource, Omnistrate automatically exposes the custom DNS hostname as an input parameter for your customers. This means your customers can provide their desired domain name directly during instance creation or update it afterward, without requiring additional configuration from the SaaS Provider. The input parameter appears in the Customer Portal and API as a configurable field on the Resource. When a customer provides a domain name, Omnistrate provisions the necessary infrastructure (L7 load balancer, TLS certificate) and returns the TXT verification record that the customer must add to their DNS configuration. ### How it works 1. **SaaS Provider** enables `customDNS` capability in the Compose Spec or Plan Spec 1. **Customer** provides their desired domain name as an input parameter when creating or updating an instance 1. **Omnistrate** provisions the L7 load balancer and generates a TXT verification record 1. **Customer** adds the TXT record and CNAME or A record to their DNS provider 1. **Omnistrate** verifies domain ownership and provisions the TLS certificate Note The custom DNS input parameter is automatically generated when you enable the `customDNS` capability. You do not need to manually define it as an API parameter. ## Setting up custom endpoint aliases - Users can register or transfer a custom domain name through a domain registrar of their choice. - Once the domain is acquired, users can update their custom resource instance endpoint with the newly acquired custom domain name. - Next, users need to configure DNS settings and create CNAME or A record to map their custom domain to the target endpoint provided by Omnistrate. - Additionally, users must add a TXT record with the "verification-" prefix to their custom domain, using the instance ID as the value to verify domain ownership. - Omnistrate facilitates secure communication over HTTPS, with certificates issued by trusted public certificate authorities (CAs), such as Google CA. After configuring DNS settings, SSL certificates, and endpoint configurations, users can validate domain ownership and initiate DNS propagation to ensure that domain mappings are applied correctly. Domain propagation may take some time to propagate globally and become accessible to users worldwide. Note Users can configure a single alias for each resource instance. Adding a new alias will replace the existing alias configuration. Omnistrate only supports TLS/SSL encrypted communication. GCP only supports A record configuration for alias mapping. AWS only supports CNAME configuration for alias mapping. # GPU Accelerator Configuration Guide ## Enabling GPU Accelerators GPU accelerator configuration enables you to specify dedicated GPU resources for your SaaS Products. This feature allows your services to leverage hardware acceleration for AI/ML workloads, high-performance computing, graphics processing, and other GPU-intensive tasks. This guide covers two ways to use GPUs with your SaaS Products: 1. **External GPU attachment** using `acceleratorConfiguration` (GCP N1 instances + Tesla GPUs only) 1. **Built-in GPU instances** across AWS, GCP (G2, A2, A3, A4 families), and Azure (modern GPUs pre-installed) ## External GPU Attachment with acceleratorConfiguration The `acceleratorConfiguration` feature is part of the `configurationOverrides` section in your service compute configuration. It provides a declarative way to specify the type and count of GPU accelerators that should be attached to your compute instances. Critical Limitations External GPU attachment via `acceleratorConfiguration` is **only supported** with: ``` - **Instance Family**: N1 instances only - **GPU Types**: Tesla series GPUs only (T4, V100, P100, P4) - **Cloud Provider**: Google Cloud Platform only All other instance families (E2, C2, M1, N2, etc.) **cannot** attach external GPUs. ``` ### Configuration Properties Each accelerator configuration requires: - **Type**: The Tesla GPU accelerator type (required) - **Count**: Number of GPU accelerators to attach (required, minimum: 1) ## Supported Accelerator Types These are the **only** GPU types that can be attached to N1 instances: - `nvidia-tesla-t4` - NVIDIA Tesla T4 (16GB GDDR6) - `nvidia-tesla-v100` - NVIDIA Tesla V100 (16GB/32GB HBM2) - `nvidia-tesla-p100` - NVIDIA Tesla P100 (16GB HBM2) - `nvidia-tesla-p4` - NVIDIA Tesla P4 (8GB GDDR5) - `nvidia-tesla-p4-vws` - NVIDIA Tesla P4 Virtual Workstation (8GB GDDR5) - `nvidia-tesla-t4-vws` - NVIDIA Tesla T4 Virtual Workstation (16GB GDDR6) - `nvidia-tesla-p100-vws` - NVIDIA Tesla P100 Virtual Workstation (16GB HBM2) ### Examples #### Basic Tesla T4 configuration (Docker Compose Style) ``` x-omnistrate-compose-spec: services: gpu-service: x-omnistrate-compute: instanceTypes: - name: n1-standard-8 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-t4" count: 1 ``` #### High-performance Tesla V100 configuration (Omnistrate Spec) ``` services: - name: gpu-v100-service compute: instanceTypes: - name: n1-standard-4 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-v100" count: 1 ``` #### Multi-GPU setup with Tesla T4 ``` services: - name: gpu-cluster compute: instanceTypes: - name: n1-standard-16 cloudProvider: gcp configurationOverrides: acceleratorConfiguration: type: "nvidia-tesla-t4" count: 4 ``` ## Built-in GPU Instances (Alternative Approach) If you need **modern GPUs** (L4, A100, H100), you should use **built-in GPU instance families** instead of `acceleratorConfiguration`. These instances come with GPUs pre-installed and offer better performance. Built-in GPU instances are available across **AWS, GCP, and Azure**. No Configuration Required Built-in GPU instances **do not** use `acceleratorConfiguration`. Simply specify the instance type and cloud provider - the GPUs are already included. ### GCP built-in GPU instance families #### G2 family - NVIDIA L4 GPUs - **GPU**: NVIDIA L4 (24GB GDDR6) - **Use Cases**: AI inference, machine learning, graphics workloads - **Instance Types**: `g2-standard-4`, `g2-standard-8`, `g2-standard-16`, etc. #### A2 family - NVIDIA A100 GPUs - **GPU**: NVIDIA A100 (40GB HBM2e) - **Use Cases**: High-performance training, large-scale inference - **Instance Types**: `a2-highgpu-1g`, `a2-highgpu-2g`, `a2-highgpu-4g`, etc. #### A3 family - NVIDIA H100 GPUs - **GPU**: NVIDIA H100 (80GB HBM3) - **Use Cases**: Advanced AI training, large language models - **Instance Types**: `a3-highgpu-8g`, `a3-megagpu-8g`, etc. #### A4 family - NVIDIA L4 GPUs - **GPU**: NVIDIA L4 (24GB GDDR6) - **Use Cases**: AI inference, video processing - **Instance Types**: `a4-highgpu-1g`, `a4-highgpu-2g`, etc. ### GCP built-in GPU examples #### Using G2 with built-in L4 GPUs ``` x-omnistrate-compose-spec: services: modern-gpu-service: x-omnistrate-compute: instanceTypes: - name: g2-standard-8 cloudProvider: gcp # No acceleratorConfiguration needed - L4 GPU is built-in ``` #### Using A2 with built-in A100 GPUs ``` services: - name: training-service compute: instanceTypes: - name: a2-highgpu-1g cloudProvider: gcp # No acceleratorConfiguration needed - A100 GPU is built-in ``` #### Using A3 with built-in H100 GPUs ``` services: - name: llm-service compute: instanceTypes: - name: a3-highgpu-8g cloudProvider: gcp # No acceleratorConfiguration needed - H100 GPUs are built-in ``` ### AWS built-in GPU instance families #### G4dn family - NVIDIA T4 GPUs - **GPU**: NVIDIA T4 (16GB GDDR6) - **Use Cases**: AI inference, graphics-intensive applications, machine learning - **Instance Types**: `g4dn.xlarge`, `g4dn.2xlarge`, `g4dn.4xlarge`, `g4dn.8xlarge`, `g4dn.12xlarge`, `g4dn.16xlarge` #### G5 family - NVIDIA A10G GPUs - **GPU**: NVIDIA A10G (24GB GDDR6) - **Use Cases**: AI inference, machine learning training, graphics workloads - **Instance Types**: `g5.xlarge`, `g5.2xlarge`, `g5.4xlarge`, `g5.8xlarge`, `g5.12xlarge`, `g5.16xlarge`, `g5.48xlarge` #### G6 family - NVIDIA L4 GPUs - **GPU**: NVIDIA L4 (24GB GDDR6) - **Use Cases**: AI inference, video transcoding - **Instance Types**: `g6.xlarge`, `g6.2xlarge`, `g6.4xlarge`, `g6.8xlarge`, `g6.12xlarge`, `g6.16xlarge`, `g6.48xlarge` #### P4d family - NVIDIA A100 GPUs - **GPU**: NVIDIA A100 (40GB HBM2e) - **Use Cases**: High-performance training, large-scale inference, HPC - **Instance Types**: `p4d.24xlarge` #### P5 family - NVIDIA H100 GPUs - **GPU**: NVIDIA H100 (80GB HBM3) - **Use Cases**: Advanced AI training, large language models, generative AI - **Instance Types**: `p5.48xlarge` ### AWS built-in GPU examples #### Using G4dn with built-in T4 GPUs ``` services: - name: inference-service compute: instanceTypes: - name: g4dn.xlarge cloudProvider: aws # T4 GPU is built-in ``` #### Using G5 with built-in A10G GPUs ``` services: - name: ml-training-service compute: instanceTypes: - name: g5.2xlarge cloudProvider: aws # A10G GPU is built-in ``` ### Azure built-in GPU instance families #### NC T4 v3 series - NVIDIA T4 GPUs - **GPU**: NVIDIA T4 (16GB GDDR6) - **Use Cases**: AI inference, machine learning, graphics - **Instance Types**: `Standard_NC4as_T4_v3`, `Standard_NC8as_T4_v3`, `Standard_NC16as_T4_v3`, `Standard_NC64as_T4_v3` #### NCv3 series - NVIDIA V100 GPUs - **GPU**: NVIDIA V100 (16GB HBM2) - **Use Cases**: Deep learning training, HPC simulations - **Instance Types**: `Standard_NC6s_v3`, `Standard_NC12s_v3`, `Standard_NC24s_v3` #### NC A100 v4 series - NVIDIA A100 GPUs - **GPU**: NVIDIA A100 (80GB HBM2e) - **Use Cases**: Large-scale AI training, inference, HPC - **Instance Types**: `Standard_NC24ads_A100_v4`, `Standard_NC48ads_A100_v4`, `Standard_NC96ads_A100_v4` #### ND H100 v5 series - NVIDIA H100 GPUs - **GPU**: NVIDIA H100 (80GB HBM3) - **Use Cases**: Advanced AI training, large language models, generative AI - **Instance Types**: `Standard_ND96isr_H100_v5` ### Azure built-in GPU examples #### Using NC T4 v3 with built-in T4 GPUs ``` services: - name: inference-service compute: instanceTypes: - name: Standard_NC4as_T4_v3 cloudProvider: azure # T4 GPU is built-in ``` #### Using NC A100 v4 with built-in A100 GPUs ``` services: - name: training-service compute: instanceTypes: - name: Standard_NC24ads_A100_v4 cloudProvider: azure # A100 GPU is built-in ``` ## Multi-Cloud GPU Configuration You can configure GPU instances across multiple cloud providers in the same Plan specification. The platform automatically selects the appropriate instance type based on the customer's deployment region and cloud provider. ``` services: - name: gpu-inference compute: instanceTypes: - name: g4dn.xlarge cloudProvider: aws - name: g2-standard-4 cloudProvider: gcp - name: Standard_NC4as_T4_v3 cloudProvider: azure ``` # GPU Slicing and Multi-Tenant GPU Configuration ## Introduction Graphics Processing Units (GPUs) have evolved far beyond their original purpose of rendering graphics. Today, they are essential for artificial intelligence (AI), machine learning (ML), and high-performance computing workloads. NVIDIA GPUs, in particular, dominate the ML landscape due to their specialized architecture designed for parallel processing, making them indispensable for matrix operations and mathematical computations that power AI-driven insights. However, GPUs are expensive resources, and maximizing their utilization is crucial for cost-effective operations. Traditional GPU allocation methods often lead to underutilization, where a single workload consumes an entire GPU but only uses a fraction of its computational capacity. This is where GPU sharing strategies become essential. ## NVIDIA GPU Sharing Strategies NVIDIA provides several approaches to enable multiple workloads to share GPU resources efficiently, including: ### 1. Time-Slicing Time-slicing divides GPU access into small time intervals, allowing different tasks to use the GPU in predefined time slices. This approach is similar to how CPUs time-slice between different processes. **How it works:** - Multiple processes share a single GPU by taking turns accessing it - The GPU scheduler allocates time slots to different workloads - No memory isolation between processes - Suitable for workloads with intermittent GPU usage patterns **Ideal use cases:** - Development and testing environments - Multiple small-scale workloads - Batch processing tasks - Real-time analytics with streaming data - Cost-sensitive environments with budget constraints ### 2. Multi-Instance GPU (MIG) MIG is available on NVIDIA A100, A30, and H100 GPUs, allowing a single physical GPU to be partitioned into multiple isolated instances. Each instance has its own dedicated memory, cache, and compute cores. **How it works:** - Physical GPU is partitioned into up to 7 separate instances - Each instance provides guaranteed performance and memory isolation - Complete fault isolation between instances - Hardware-level partitioning ensures predictable performance **Ideal use cases:** - Multi-tenant environments requiring strict isolation - Cloud service providers offering GPU-as-a-Service - Workloads requiring guaranteed performance levels - Environments with strict SLA requirements ## Time-Slicing vs MIG: Detailed Comparison | Aspect | Time-Slicing | Multi-Instance GPU (MIG) | | -------------------------- | ------------------------------------ | -------------------------------------- | | **Memory Isolation** | No isolation - shared memory space | Complete memory isolation per instance | | **Fault Isolation** | No isolation - one crash affects all | Complete fault isolation | | **Performance Guarantees** | No guarantees - best effort sharing | Guaranteed performance per instance | | **GPU Support** | All NVIDIA GPUs | Limited to A100, A30, H100 | | **Resource Overhead** | Minimal overhead | Some overhead due to partitioning | | **Use Case** | Resource optimization | Strong Isolation and SLA requirements | ## Configuring GPU Slicing DIY Enabling and managing GPU slicing infrastructure is a complex undertaking that involves multiple layers of technology, careful resource planning, and ongoing operational overhead. While platforms like Omnistrate abstract much of this complexity, understanding the underlying challenges helps appreciate the engineering effort required to make GPU sharing work effectively. ### Infrastructure Complexity Overview GPU slicing requires orchestrating several complex systems that must work together seamlessly: #### 1. **Hardware-Level Considerations** - **GPU Architecture Compatibility**: Not all GPUs support the same sharing mechanisms. NVIDIA's MIG is only available on A100, A30, and H100 GPUs, while time-slicing works across different GPU generations but with varying performance characteristics #### 2. **Kernel and Driver Stack Complexity** - **NVIDIA Driver Management**: Requires specific driver versions that support sharing features, with complex upgrade paths that can break existing workloads - **CUDA Runtime Coordination**: Managing CUDA contexts across multiple processes requires sophisticated scheduling and memory management #### 3. **Container Orchestration Challenges** - **Device Plugin Architecture**: Implementing and maintaining custom Kubernetes device plugins that can advertise virtual GPU resources accurately - **Resource Scheduling Complexity**: The Kubernetes scheduler must understand GPU topology, memory constraints, and performance characteristics to make optimal placement decisions - **Namespace Isolation**: Ensuring proper isolation between different tenants while maintaining GPU access ### Operational Management Complexity #### **Capacity Planning and Resource Allocation** - **Workload Characterization**: Understanding the GPU usage patterns of different workloads to optimize sharing ratios - **Performance Modeling**: Predicting how different combinations of workloads will perform when sharing GPU resources - **Cost Optimization**: Balancing the cost of GPU instances against the performance impact of sharing - **Scaling Strategies**: Determining when to scale horizontally (more GPU instances) vs. vertically (more sharing on existing GPUs) #### **Lifecycle Management** - **Rolling Updates**: Updating GPU drivers, CUDA versions, or container runtimes without disrupting running workloads - **Workload Migration**: Moving workloads between different GPU instances during maintenance or optimization - **Disaster Recovery**: Implementing backup and recovery strategies for stateful GPU workloads - **Version Compatibility**: Managing compatibility matrices between CUDA versions, driver versions, and application requirements ### Why Managed Solutions Matter The complexity outlined above explains why managed platforms like Omnistrate provide significant value: - **Abstraction of Complexity**: Hiding the intricate details of GPU driver management, device plugin configuration, and monitoring setup - **Tested Configurations**: Providing pre-validated combinations of hardware, software, and configuration that work reliably together - **Automated Operations**: Handling routine maintenance, updates, and optimization tasks automatically - **Expert Support**: Access to specialists who understand the nuances of GPU sharing infrastructure ## Configuring GPU Slicing in Omnistrate Omnistrate provides built-in support for GPU slicing through the `multiTenantGpu` feature, making it easy to deploy services that efficiently share GPU resources across multiple tenants. ### Configuration Overview To enable GPU slicing in your Omnistrate service, you need to add the `x-internal-integrations` section to your `omnistrate-compose.yaml` file: ``` x-internal-integrations: multiTenantGpu: instanceType: g4dn.xlarge # instance type to be used for GPU slicing timeSlicingReplicas: 2 # number of replicas to be used for time slicing migProfile: 1g.5gb # optional: MIG profile for A100/H100 GPUs ``` ### Configuration Parameters #### `instanceType` Specifies the EC2 instance type that will host the GPU slicing functionality. Common GPU-enabled instance types include: - **g4dn.xlarge**: 1 NVIDIA T4 GPU, 4 vCPUs, 16 GB RAM - Cost-effective for inference workloads - **g4dn.2xlarge**: 1 NVIDIA T4 GPU, 8 vCPUs, 32 GB RAM - Balanced compute and memory - **p3.2xlarge**: 1 NVIDIA V100 GPU, 8 vCPUs, 61 GB RAM - High-performance training - **p4d.24xlarge**: 8 NVIDIA A100 GPUs, 96 vCPUs, 1152 GB RAM - Multi-GPU training #### `timeSlicingReplicas` Defines how many virtual GPU replicas will be created from each physical GPU. This determines how many concurrent workloads can share a single GPU. - **Value of 2**: Each physical GPU appears as 2 virtual GPUs (50% allocation per workload) - **Value of 4**: Each physical GPU appears as 4 virtual GPUs (25% allocation per workload) - **Value of 8**: Each physical GPU appears as 8 virtual GPUs (12.5% allocation per workload) #### `migProfile` (Optional) Specifies the Multi-Instance GPU (MIG) profile to use when the instance type supports MIG (A100, A30, H100 GPUs). This parameter enables hardware-level GPU partitioning with guaranteed isolation and performance. Note MIG and Time-Slicing be combined to create a multi-layered GPU sharing strategy. You can use both `migProfile` and `timeSlicingReplicas` together to further subdivide MIG instances with time-slicing. ### Complete Example Configuration Here are complete examples of GPU-sliced service configurations: #### Time-Slicing Configuration Example ``` version: '3.9' x-omnistrate-service-plan: name: 'gpu-slicing-example hosted tier' tenancyType: 'OMNISTRATE_MULTI_TENANCY' x-internal-integrations: multiTenantGpu: instanceType: g4dn.xlarge # instance type to be used for GPU slicing timeSlicingReplicas: 2 # number of replicas to be used for time slicing services: gpuinfo: image: ghcr.io/omnistrate-community/gpu-slicing-example:0.0.3 ports: - 5000:5000 platform: linux/amd64 deploy: resources: limits: cpus: '1' memory: 100M reservations: cpus: '100m' memory: 50M x-omnistrate-capabilities: autoscaling: maxReplicas: 3 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 serverlessConfiguration: targetPort: 5000 enableAutoStop: true minimumNodesInPool: 1 ``` #### MIG Configuration Example ``` version: '3.9' x-omnistrate-service-plan: name: 'gpu-mig-example hosted tier' tenancyType: 'OMNISTRATE_MULTI_TENANCY' x-internal-integrations: multiTenantGpu: instanceType: p4d.24xlarge # A100 GPU instance supporting MIG migProfile: 1g.5gb # MIG profile for hardware-level isolation services: gpuinfo: image: ghcr.io/omnistrate-community/gpu-slicing-example:0.0.3 ports: - 5000:5000 platform: linux/amd64 deploy: resources: limits: cpus: '2' memory: 500M reservations: cpus: '500m' memory: 250M x-omnistrate-capabilities: autoscaling: maxReplicas: 5 minReplicas: 1 idleMinutesBeforeScalingDown: 2 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 3 overUtilizedThreshold: 80 serverlessConfiguration: targetPort: 5000 enableAutoStop: true minimumNodesInPool: 1 ``` ### How GPU Slicing Works in Omnistrate 1. **Infrastructure Provisioning**: Omnistrate provisions the specified GPU-enabled EC2 instance 1. **NVIDIA Device Plugin**: Automatically installs and configures the NVIDIA Kubernetes device plugin 1. **Time-Slicing Configuration**: Configures the device plugin with the specified number of replicas 1. **Resource Advertisement**: The Kubernetes scheduler sees multiple virtual GPUs instead of one physical GPU 1. **Workload Scheduling**: Multiple pods can be scheduled to share the same physical GPU 1. **Automatic Scaling**: Omnistrate can automatically scale the number of replicas based on demand ### Best Practices #### Choosing the Right Instance Type - **For inference workloads**: Use g4dn instances with T4 GPUs for cost-effectiveness - **For training workloads**: Use p3 instances with V100 GPUs for better performance - **For large-scale training**: Consider p4d instances with A100 GPUs and MIG support #### Setting Time-Slicing Replicas - **Start conservative**: Begin with 2-4 replicas and monitor performance - **Monitor GPU utilization**: Use tools like `nvidia-smi` to track actual GPU usage - **Consider workload characteristics**: CPU-bound tasks can share more aggressively than GPU-intensive ones - **Account for memory usage**: Ensure total GPU memory usage doesn't exceed physical limits #### Resource Management - Set appropriate CPU and memory limits for your containers - Use Omnistrate's autoscaling capabilities to handle varying demand - Monitor performance metrics to optimize replica counts - Consider using different configurations for development vs production environments ## Example Use Cases ### AI/ML Model Serving Deploy multiple model inference endpoints that share GPU resources efficiently: - Each model gets dedicated time slices for inference - Cost-effective serving of multiple models - Automatic scaling based on request volume ### Development and Testing Enable multiple developers to share GPU resources: - Each developer gets access to GPU acceleration - Reduced infrastructure costs for development teams - Isolated development environments ### Batch Processing Process multiple data pipelines concurrently: - Different batch jobs share GPU resources - Improved throughput for data processing workflows - Cost optimization for periodic workloads ## Conclusion GPU slicing with Omnistrate provides a powerful way to maximize GPU utilization while minimizing costs. By leveraging NVIDIA's time-slicing technology through simple configuration parameters, you can enable multiple workloads to efficiently share expensive GPU resources. Omnistrate's built-in GPU slicing support makes it easy to implement either approach, allowing you to focus on your application logic while the platform handles the complex GPU resource management automatically. # Local NVMe Storage Many instance types ship with NVMe disks physically attached to the host. These are far faster than network-attached volumes (EBS, Persistent Disk, Azure Disk) and are the right backing store for scratch space, caches, spill files, and analytical workloads that re-read large working sets. You reach them the same way on every cloud — a claim on the `omnistrate-local-nvme` storage class — but how much you get differs: | Cloud | How much local NVMe you get | | --------- | ---------------------------------------------------- | | **AWS** | Fixed by the instance type. Pick one that has it. | | **Azure** | Fixed by the instance type. Pick one that has it. | | **GCP** | A count you choose, set in `configurationOverrides`. | Local NVMe is ephemeral Data on local NVMe does not survive the node going away. It is lost on node deallocation, reimage, scaling down, and instance replacement. Use it for data you can rebuild — never as the only copy of anything you need to keep. For durable data, use a [persistent volume](https://docs.omnistrate.com/infra-guides/persistent-volumes/index.md). ## Using Local NVMe How you reach it depends on how your resource is defined. ### Compose-based resources Nothing to configure. When the instance type has local NVMe, your container gets a `/cache` directory backed by it automatically on AWS and Azure. On GCP, also set [`ephemeralStorageLocalSsdConfig`](#gcp-choosing-how-many-local-ssds) — local SSD is an attached resource there, so the count has to be requested. Use `/cache` for scratch space, spill files and caches. Your persistent volumes stay on network-attached storage, which is what you want for anything that has to survive the node. Persistent volumes cannot be placed on local NVMe from a Compose spec `x-omnistrate-storage` describes network-attached disks — EBS, Persistent Disk, Azure Disk — so `instanceStorageType` has no local-NVMe option. `/cache` is the supported way to use local NVMe from a Compose spec. ### Helm chart, Kustomize and Operator resources Omnistrate does not control your pod spec here, so there is no automatic `/cache`. Ask for local NVMe with a `PersistentVolumeClaim` on the `omnistrate-local-nvme` storage class and mount it wherever your image expects it: ``` volumeClaimTemplates: - metadata: name: cache spec: accessModes: [ReadWriteOnce] storageClassName: omnistrate-local-nvme resources: requests: storage: 100Gi ``` ``` volumeMounts: - name: cache mountPath: /cache ``` Mounting at `/cache` keeps an image portable across deployment types: it sees the same path whether it runs from a Compose spec, a Helm chart, a Kustomize overlay or an Operator. Any other path works just as well if your image expects something else. The claim is identical on every cloud — Omnistrate prepares the disks and decides where the bytes live, so it does not change when you move between AWS, Azure and GCP. On GCP, set [`ephemeralStorageLocalSsdConfig`](#gcp-choosing-how-many-local-ssds) as well, since local SSD is an attached resource there and the count has to be requested. Nothing else is required: no `hostPath`, no `mountPropagation`, no init container, no privileged pod, no node selector. The claim is only scheduled onto nodes that actually have local NVMe, so it cannot quietly land on a node without it and give you OS-disk performance. Because the volume is a directory on one specific node, the claim stays `Pending` until a pod using it is scheduled — `waiting for first consumer` in its events is normal — and the pod is then pinned to the node holding its data. Sizing is advisory `resources.requests.storage` is not enforced. The volume is a directory on the node's array, so it can grow until the array is full and is not isolated from other volumes on the same node. Size your node pool for the total you need. ### What Omnistrate does for you At node bootstrap it finds the local NVMe devices, ignoring the OS disk and any attached data disks, combines them into a single RAID 0 array when there is more than one, and formats and mounts it. `/cache` and any local-NVMe volume claims are served from that array. ### Choosing an Instance Type with Local NVMe **AWS** — families whose name carries a `d`, plus the storage-optimized families: - `m6id`, `m7gd`, `c6id`, `c7gd`, `r6id`, `r7gd` and similar `d`-suffixed variants - `i3`, `i3en`, `i4i`, `im4gn`, `is4gen` ``` aws ec2 describe-instance-types --instance-types m6id.8xlarge \ --query 'InstanceTypes[].InstanceStorageInfo' ``` A non-empty result with `"NvmeSupport": "required"` means it has local NVMe. **Azure** — the v6/v7 `d` families, the storage-optimized L-series, and several HPC and GPU families: - `Ddsv6`, `Edsv6`, `Ddsv7`, `Edsv7`, `Eadsv7`, `Dldsv6`, `Dpdsv6`, `Epdsv6`, `Fadsv7` - `Lsv2`, `Lsv3`, `Lasv3`, `Lsv4`, `Lasv4`, `Laosv4` - `FXmdsv2`, `HBv4`, `HX` - NVMe-backed GPU families such as `NCadsH100v5` and `NCADSA100v4` ``` az vm list-skus --location eastus2 --size Standard_E32ds_v6 -o json \ | jq -r '.[0].capabilities[] | select(.name | startswith("Nvme"))' ``` A non-zero `NvmeDiskSizeInMiB` means it has local NVMe. Divide it by `NvmeSizePerDiskInMiB` for the number of disks. Azure: a temp disk is not the same thing Azure's `MaxResourceVolumeMB` describes the legacy temp/resource disk, a different device. Some SKUs have both, some only one. Only `NvmeDiskSizeInMiB` tells you about local NVMe. ## GCP: Choosing How Many Local SSDs On GCP, local SSD is an attached resource: for many machine families you choose how many disks to attach. That count is set with `configurationOverrides` on the instance type. You still consume the result through the same storage class. Contact support to enable Local SSD on GCP is gated. Email [support@omnistrate.com](mailto:support@omnistrate.com) to have it enabled for your organization. There are two fields, for the two ways GCP exposes local SSD. ### ephemeralStorageLocalSsdConfig Backs the node's **ephemeral storage** with local SSD, so `emptyDir` volumes and container image layers land on it. This is the closest equivalent to the AWS and Azure behavior above. | Property | Description | | ---------------- | --------------------------------------------------- | | `localSsdCount` | Number of local SSDs to back ephemeral storage with | | `dataCacheCount` | Number of local SSDs to reserve for data caching | At least one of the two must be set, and any value given must be greater than zero. ``` x-omnistrate-compose-spec: services: warehouse: x-omnistrate-compute: instanceTypes: - name: n2-standard-8 cloudProvider: gcp configurationOverrides: ephemeralStorageLocalSsdConfig: localSsdCount: 2 ``` ### localNvmeSsdBlockConfig Attaches local SSDs as **raw block devices**, left unformatted for the workload to manage itself. Use this when your application wants to own the device directly. | Property | Description | | --------------- | --------------------------------------------- | | `localSsdCount` | Number of raw-block local NVMe SSDs to attach | `localSsdCount` is required and must be greater than zero. ``` services: - name: warehouse compute: instanceTypes: - name: n2-standard-8 cloudProvider: gcp configurationOverrides: localNvmeSsdBlockConfig: localSsdCount: 1 ``` The number of local SSDs you can attach depends on the machine type and zone. See [Google's local SSD documentation](https://cloud.google.com/compute/docs/disks/local-ssd) for the limits. ### These Fields Are GCP-Only Both fields are rejected at spec import for any other cloud provider: ``` ephemeral storage local SSD configuration is not supported for AWS local NVMe SSD block configuration is not supported for Azure ``` This is intentional, not a gap. On AWS and Azure the amount of local NVMe is fixed by the instance type and there is no count to choose — see [Using Local NVMe](#using-local-nvme) above. ## Verifying It Works From inside a pod, check where its volume actually lives: ``` kubectl exec -- df -h /path/to/volume ``` A volume on local NVMe shows a device such as `/dev/nvme1n1` or `/dev/md/...`. If it shows the root device instead (`/dev/root` on Azure), the instance type most likely has no local NVMe — confirm with the `aws` or `az` command in [Choosing an Instance Type with Local NVMe](#choosing-an-instance-type-with-local-nvme). If a claim stays `Pending`, check its events: ``` kubectl describe pvc ``` `waiting for first consumer` is normal — the volume is a directory on one specific node, so it is not created until a pod using it has been scheduled. ## Related Guides - [Compute Management](https://docs.omnistrate.com/infra-guides/compute/index.md) — selecting instance types per cloud provider - [Persistent Volumes](https://docs.omnistrate.com/infra-guides/persistent-volumes/index.md) — durable storage that survives node replacement - [Storage Classes](https://docs.omnistrate.com/infra-guides/storage-classes/index.md) — pre-provisioned classes for network-attached storage # Infra Guide Overview Omnistrate offer turn key solution to manage your infrastructure. You can use Omnistrate-native Infrastructure management. Here are some of the advantages: - Tenant-aware: manage per-tenant Infrastructure - ACID-compliant: make sure all Infra operations are fully ACID-compliant and not leave the leaked resources in between. - Versioned: all changes are versioned - Multi-cloud: abstracts away the complexities of the cloud without losing control - Day-2 Ops: automates infrastructure operations, not just provisioning You can read more [here](https://blog.omnistrate.com/posts/58). Alternatively, you can bring your Terraform or equivalent and incorporate with other building blocks. Here are some of the core infra capabilities: - Endpoint aliases to provide custom domain name for your Resource. To learn more, please see [here](https://docs.omnistrate.com/infra-guides/endpoint-aliases/index.md) - Configure load balancers to load balance across multiple nodes. To learn more, see [here](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-load-balancer) - Shared File System support to save data in a persistent volume and share it across different pods. To learn more, please see [here](https://docs.omnistrate.com/infra-guides/shared-file-system/index.md) - Blob Storage to define blob storage(s) and their mount paths across Resources. To learn more, please see [here](https://docs.omnistrate.com/infra-guides/blob-storage/index.md) - GPU Accelerator to specify dedicated GPU resources for your SaaS Products. To learn more, please see [here](https://docs.omnistrate.com/infra-guides/gpu-accelerator-configuration/index.md) - Custom Deployment Cell Placement to allow you to control how many deployments can be co-located on the same host cluster (deployment cell / Kubernetes cluster). To learn more, please see [here](https://docs.omnistrate.com/infra-guides/custom-deployment-cell-placement/index.md) If you don't see your favorite capability above, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com). We would love to understand your use-case and prioritize the support. # Persistent volumes To define a persistent volumes that can survive service restarts and nodes replacements you can use `x-omnistrate-storage` tag. Using this tag the defined volumes in the compose spec can become persistent volumes on the deployed instance. Omnistrate will take care of provisioning and managing the volumes. You can optionally define the volume, and then define the mount path and properties for the volume. ``` services: database: image: postgres volumes: - source: pg_master_data target: /var/lib/postgresql/data type: volume x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP3 instanceStorageSizeGi: 10 instanceStorageIOPS: 3000 instanceStorageThroughputMiBps: 125 gcp: instanceStorageType: GCP::PD_BALANCED instanceStorageSizeGi: 100 azure: instanceStorageType: AZURE::PREMIUM_SSD instanceStorageSizeGi: 100 volumes: pg_master_data: driver: local ``` You can customize the following: - Storage type from block device to blobs or both - Size of the volume - Storage IOPS - Storage throughput Note This type of persistent volume are not shared across Resources. If you need to share the volume data across resources look a at configuring Shared File System or Blob Storage instead # Shared File System ## Why using a shared file system In modern SaaS application deployments on Kubernetes, you often need to save data in a persistent volume and share it across different pods. For example, when setting up an AI model training system, multiple workers might need access to the same dataset stored in a persistent volume. This is where a shared file system is a good fit. Additionally, if your data volume is growing and you need to scale the volume size dynamically without downtime, a shared file system is an excellent solution for you. Omnistrate allows your Resources to share data within the same deployment cluster by utilizing a shared file system. It provides the capability to easily share data across Resources and to scale your underlying storage without any downtime. ## Configuring shared file system To configure a shared file system using compose specification you can define the shared volume at the root level and then define the volume and mount for each resource. Here is the example on how to define and mount a shared volume: ``` services: serviceA: volumes: - source: file_system_data target: /var/data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::EFS gcp: clusterStorageType: GCP::FILESTORE azure: clusterStorageType: AZURE::FILE_SHARE nebius: clusterStorageType: NEBIUS::FILESYSTEM_NETWORK_SSD volumes: file_system_data: driver: sharedFileSystem driver_opts: # efs configuration efsThroughputMode: provisioned efsPerformanceMode: generalPurpose efsProvisionedThroughputInMibps: 100 # filestore configuration filestoreCapacityGi: 1024 filestoreTier: "BASIC_HDD" # azure fileshare configuration fileshareQuotaGi: 1024 fileshareTier: Premium fileshareRedundancy: LRS # Nebius shared filesystem configuration nebiusFilesystemType: NETWORK_SSD nebiusFilesystemSizeGi: 1024 ``` **service volume properties**: - `source` is the name of the volume you defined in the previous step. - `target` is the path where you want to mount the volume. - `type` specifies the type of the volume; in this case, it is `volume`. - `x-omnistrate-storage` defines the type of storage for each cloud provider. **volume properties**: - `file_system_data` is the name of the volume, you can opt for a name of your choice. - `driver: sharedFileSystem` is the driver type required to enable shared file system functionality. - `driver_opts` defines various options to customize the shared file system configuration for different cloud providers: - `efsThroughputMode`: Specifies the throughput mode for AWS EFS. Can be set to `bursting` or `provisioned`. Use `provisioned` when you need consistent throughput performance. See [AWS EFS Throughput Modes](https://docs.aws.amazon.com/efs/latest/ug/performance.html#throughput-modes) for details. - `efsPerformanceMode`: Defines the performance mode for AWS EFS. Options are `generalPurpose` (for latency-sensitive workloads) or `maxIO` (for higher aggregate throughput and operations per second). See [AWS EFS Performance Modes](https://docs.aws.amazon.com/efs/latest/ug/performance.html#performance-modes) for more information. - `efsProvisionedThroughputInMibps`: Sets the provisioned throughput in MiB/s for AWS EFS when using `provisioned` throughput mode. This value determines the baseline throughput performance. See [AWS EFS Provisioned Throughput](https://docs.aws.amazon.com/efs/latest/ug/performance.html#provisioned-throughput) for details. - `filestoreCapacityGi`: Specifies the storage capacity for GCP Filestore in GiB. Minimum capacity requirements vary by tier: `BASIC_HDD`, `ENTERPRISE`, `REGIONAL`, and `ZONAL` require at least 1024 and `BASIC_SSD` requires at least 2560. For the most up-to-date details, see the [GCP Filestore documentation](https://cloud.google.com/filestore/docs/tiers#comparison). - `filestoreTier`: Defines the service tier for GCP Filestore. Options include `BASIC_HDD`, `BASIC_SSD`, `ENTERPRISE`, `REGIONAL`, or `ZONAL` depending on your performance and capacity requirements. - `fileshareQuotaGi`: Specifies the quota for Azure File Share in GiB. The quota must be between 1 GiB and 102400 GiB. - `fileshareTier`: Defines the performance tier for Azure File Share. Options include `Standard` or `Premium`. - `fileshareRedundancy`: Specifies the redundancy option for Azure File Share. Options include `LRS` (Locally Redundant Storage), `ZRS` (Zone-Redundant Storage), `GRS` (Geo-Redundant Storage), or `GZRS` (Geo-Zone-Redundant Storage) for `Standard` tier and `LRS` (Locally Redundant Storage) or `ZRS` (Zone-Redundant Storage) for `Premium` tier. - `nebiusFilesystemType`: Defines the Nebius shared filesystem backend type. Supported values are `NETWORK_SSD`, `NETWORK_HDD`, `WEKA`, and `VAST`. - `nebiusFilesystemSizeGi`: Specifies the Nebius shared filesystem size in GiB. For Nebius, the `clusterStorageType` in `x-omnistrate-storage` must match the backend selected in `driver_opts`: - `NETWORK_SSD` -> `NEBIUS::FILESYSTEM_NETWORK_SSD` - `NETWORK_HDD` -> `NEBIUS::FILESYSTEM_NETWORK_HDD` - `WEKA` -> `NEBIUS::FILESYSTEM_WEKA` - `VAST` -> `NEBIUS::FILESYSTEM_VAST` Nebius shared file systems follow the same control-plane pattern as the other clouds: the backend size and backend type live on the shared volume's `driver_opts`, while the per-resource mount only references the matching `clusterStorageType`. Omnistrate rejects Nebius shared file systems for `OMNISTRATE_MULTI_TENANCY` plans because Nebius filesystem attachments are realized at the nodegroup level. This configuration will provision a new Shared File System for each instance deployed. ## Mounting a shared file volume in multiple resources The same volume can be mounted in different pods: ``` services: serviceA: volumes: - source: file_system_data target: /var/data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::EFS gcp: clusterStorageType: GCP::FILESTORE azure: clusterStorageType: AZURE::FILE_SHARE nebius: clusterStorageType: NEBIUS::FILESYSTEM_NETWORK_SSD serviceB: volumes: - source: file_system_data target: /var/data type: volume x-omnistrate-storage: aws: clusterStorageType: AWS::EFS gcp: clusterStorageType: GCP::FILESTORE azure: clusterStorageType: AZURE::FILE_SHARE nebius: clusterStorageType: NEBIUS::FILESYSTEM_NETWORK_SSD volumes: file_system_data: driver: sharedFileSystem driver_opts: # efs configuration efsThroughputMode: provisioned efsPerformanceMode: generalPurpose efsProvisionedThroughputInMibps: 100 # filestore configuration filestoreCapacityGi: 1024 filestoreTier: "BASIC_HDD" # azure fileshare configuration fileshareQuotaGi: 1024 fileshareTier: Premium fileshareRedundancy: LRS # Nebius shared filesystem configuration nebiusFilesystemType: NETWORK_SSD nebiusFilesystemSizeGi: 1024 ``` # Storage Classes for Helm Charts, Operators and Kustomize ## Storage Classes Storage classes in Kubernetes define the types of storage available for dynamic provisioning of persistent volumes. When using Helm charts, Operators or Kustomize with Omnistrate, you can leverage dynamic provisioning to create volumes of different types based on your application's requirements. Omnistrate pre-provisions a defined set of storage classes for persistent volumes across different cloud providers, ensuring consistent storage options and optimal performance for your Helm deployments. When using Compose specifications using storage classes directly is not required. Looking for node-local NVMe? Every storage class below is network-attached (EBS, Persistent Disk, Azure Disk/Files, blob). Node-local NVMe is served by a separate class, `omnistrate-local-nvme`, on instance types that have it. See [Local NVMe Storage](https://docs.omnistrate.com/infra-guides/local-nvme-storage/index.md). ## Pre-Provisioned Storage Classes Omnistrate automatically provisions the following storage classes across different cloud providers: ### AWS Storage Classes | Storage Class | Type | Description | Use Cases | | --------------- | ----------- | --------------------------------------------------------- | ----------------------------------------------------------- | | `gp3` | AWS EBS GP3 | General Purpose SSD with configurable IOPS and throughput | General workloads, databases | | `gp3-immediate` | AWS EBS GP3 | General Purpose SSD with configurable IOPS and throughput | General workloads, databases | | `gp2` | AWS EBS GP2 | General Purpose SSD with baseline performance | Legacy applications, development | | `io2` | AWS EBS IO2 | Latest generation Provisioned IOPS SSD | Mission-critical applications requiring highest performance | | `io2-immediate` | AWS EBS IO2 | Latest generation Provisioned IOPS SSD | Mission-critical applications requiring highest performance | **Storage Class Configuration:** - **Volume Binding Mode**: `WaitForFirstConsumer` (default), `Immediate` (for `-immediate` variants) - **Reclaim Policy**: `Delete` - volumes are automatically deleted when deployment instances are deleted - **Volume Expansion**: Enabled for all AWS storage classes - **Encryption**: All volumes are encrypted - **Default**: `gp3` For Helm or Kustomize workloads on AWS EKS, you can create a custom EBS `StorageClass` that encrypts volumes with a customer-managed AWS KMS key. The AWS KMS key must have the tag `omnistrate.com/customer-managed-kms=true`. Note Existing customer AWS accounts must update or rerun the latest generated CloudFormation stack before using the custom storage class. Create the storage class in each deployment cell where workloads will use it: ``` apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: gp3-customer-managed-kms provisioner: ebs.csi.aws.com parameters: type: gp3 encrypted: "true" kmsKeyId: arn:aws:kms:us-east-1:123456789012:key/12345678-1234-1234-1234-123456789012 volumeBindingMode: WaitForFirstConsumer reclaimPolicy: Delete allowVolumeExpansion: true ``` Use the custom storage class from Helm values: ``` persistence: enabled: true storageClass: "gp3-customer-managed-kms" size: 100Gi ``` Use the same storage class from a Kustomize-managed `PersistentVolumeClaim`: ``` apiVersion: v1 kind: PersistentVolumeClaim metadata: name: data spec: accessModes: - ReadWriteOnce storageClassName: gp3-customer-managed-kms resources: requests: storage: 100Gi ``` ### GCP Storage Classes | Storage Class | Type | Description | Use Cases | | ------------------------ | ------------- | ---------------------------------------------- | ---------------------------------------- | | `standard-rwo` | GCP Standard | Standard ReadWriteOnce storage | Development, testing, general workloads | | `standard-rwo-immediate` | GCP Standard | Standard ReadWriteOnce storage | Development, testing, general workloads | | `pd-balanced` | GCP Balanced | Balanced persistent disk with good performance | Production workloads, databases | | `pd-balanced-immediate` | GCP Balanced | Balanced persistent disk with good performance | Production workloads, databases | | `pd-ssd` | GCP SSD | High-performance SSD persistent disk | High-performance applications, databases | | `pd-ssd-immediate` | GCP SSD | High-performance SSD persistent disk | High-performance applications, databases | | `pd-standard` | GCP Standard | Standard persistent disk | General workloads, development | | `pd-standard-immediate` | GCP Standard | Standard persistent disk | General workloads, development | | `pd-extreme` | GCP Extreme | Highest performance persistent disk | Mission-critical, high-IOPS applications | | `pd-extreme-immediate` | GCP Extreme | Highest performance persistent disk | Mission-critical, high-IOPS applications | | `hyperdisk` | GCP Hyperdisk | Ultra-high performance storage | Extreme performance requirements | | `hyperdisk-immediate` | GCP Hyperdisk | Ultra-high performance storage | Extreme performance requirements | **Storage Class Configuration:** - **Volume Binding Mode**: `WaitForFirstConsumer` (default), `Immediate` (for `-immediate` variants) - **Reclaim Policy**: `Delete` - volumes are automatically deleted when deployment instances are deleted - **Volume Expansion**: Enabled for all GCP storage classes - **Encryption**: All volumes are encrypted - **Default**: `pd-balanced` ### Azure Storage Classes | Storage Class | Type | Description | Use Cases | | ------------------------------- | ------------------------ | ------------------------------------- | ------------------------------------------- | | `managed-premium` | Azure Premium LRS | Premium locally redundant SSD storage | Production databases, high-performance apps | | `managed-premium-immediate` | Azure Premium LRS | Premium locally redundant SSD storage | Production databases, high-performance apps | | `managed-standard` | Azure Standard LRS | Standard locally redundant storage | Development, testing, general workloads | | `managed-standard-immediate` | Azure Standard LRS | Standard locally redundant storage | Development, testing, general workloads | | `managed-ultra` | Azure Ultra SSD LRS | Ultra-high performance SSD storage | Mission-critical, high-IOPS applications | | `managed-ultra-immediate` | Azure Ultra SSD LRS | Ultra-high performance SSD storage | Mission-critical, high-IOPS applications | | `azurefiles-standard` | Azure Files Standard LRS | Standard file storage | Shared file systems, development | | `azurefiles-standard-immediate` | Azure Files Standard LRS | Standard file storage | Shared file systems, development | | `azurefiles-premium` | Azure Files Premium LRS | Premium file storage | High-performance shared file systems | | `azurefiles-premium-immediate` | Azure Files Premium LRS | Premium file storage | High-performance shared file systems | **Storage Class Configuration:** - **Volume Binding Mode**: `WaitForFirstConsumer` (default), `Immediate` (for `-immediate` variants) - **Reclaim Policy**: `Delete` - volumes are automatically deleted when deployment instances are deleted - **Volume Expansion**: Enabled for `managed-premium` storage classes only - **Encryption**: All volumes are encrypted - **Default**: `managed-premium` ### Cloud Agnostic Storage Class Omnistrate configures default storage classes for each cloud, if you don't have specific performance requirements you can rely on the default storage class for simplicity. ## Using Storage classes on your Helm When deploying Helm charts through Omnistrate, you can specify storage classes in your Helm values to dynamically provision persistent volumes. The storage class determines the type and performance characteristics of the underlying storage. ### Example of using storage classes on you Helm Chart ``` # values.yaml persistence: enabled: true storageClass: "gp3" # AWS GP3 storage size: 100Gi accessMode: ReadWriteOnce ``` ``` # templates/pvc.yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: name: {{ include "myapp.fullname" . }}-data spec: accessModes: - ReadWriteOnce resources: requests: storage: {{ .Values.persistence.size }} {{- if .Values.persistence.storageClass }} storageClassName: {{ .Values.persistence.storageClass }} {{- end }} ``` ### Multi-Cloud Storage Configuration For applications that need to work across different cloud providers, you can use [layered config](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) logic in your Helm config values in Omnistrate. ## Creating Custom Storage Classes If the pre-provisioned storage classes don't meet your specific requirements, you can create your own storage classes using the pre-installed drivers that Omnistrate provides for each cloud provider. Note Custom storage classes must be created in each deployment cell where you plan to use them. Omnistrate automatically manages the lifecycle of pre-provisioned storage classes, but custom storage classes require manual management across your deployment cells. ### Available Pre-Installed Drivers Omnistrate pre-installs the following storage drivers in each deployment cell: #### Block Storage Drivers **AWS:** - `ebs.csi.aws.com` - AWS EBS CSI driver for block storage **GCP:** - `pd.csi.storage.gke.io` - GCP Persistent Disk CSI driver for block storage **Azure:** - `disk.csi.azure.com` - Azure Disk CSI driver for block storage #### File Storage Drivers **AWS:** - `efs.csi.aws.com` - AWS EFS CSI driver for shared file systems **GCP:** - `filestore.csi.storage.gke.io` - GCP Filestore CSI driver for shared file systems **Azure:** - `file.csi.azure.com` - Azure Files CSI driver for shared file systems #### Blob Storage Drivers **AWS:** - `s3.csi.aws.com` - AWS S3 CSI driver for object storage **GCP:** - `gcs.csi.storage.gke.io` - GCP Cloud Storage CSI driver for object storage **Azure:** - `blob.csi.azure.com` - Azure Blob Storage CSI driver for object storage ### Custom Storage Class Examples #### Custom Block Storage Class ``` apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: custom-high-iops provisioner: ebs.csi.aws.com parameters: type: io2 iops: "20000" throughput: "1000" encrypted: "true" volumeBindingMode: WaitForFirstConsumer reclaimPolicy: Delete allowVolumeExpansion: true ``` #### Custom File Storage Class ``` apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: custom-efs-performance provisioner: efs.csi.aws.com parameters: provisioningMode: efs-ap fileSystemId: fs-12345678 directoryPerms: "0755" performanceMode: maxIO throughputMode: provisioned provisionedThroughputInMibps: "500" volumeBindingMode: Immediate reclaimPolicy: Delete ``` #### Custom Blob Storage Class ``` apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: custom-s3-storage provisioner: s3.csi.aws.com parameters: mounter: geesefs bucketName: my-custom-bucket region: us-west-2 storageClass: STANDARD_IA volumeBindingMode: Immediate reclaimPolicy: Delete ``` ## Monitoring and Troubleshooting ### Check Storage Class Availability ``` # List available storage classes kubectl get storageclass # Describe a specific storage class kubectl describe storageclass gp3 ``` ### Monitor PVC Status ``` # Check PVC status kubectl get pvc # Describe PVC for troubleshooting kubectl describe pvc my-app-data ``` ### Common Issues 1. **PVC Stuck in Pending**: Verify the storage class exists and there is sufficient quota in the deployment cell 1. **Volume Mount Failures**: Ensure the storage class supports the requested access mode 1. **Performance Issues**: Consider upgrading to a higher-performance storage class or adjusting IOPS settings Note Omnistrate automatically handles storage class provisioning and configuration. If you encounter persistent issues, contact support for assistance with your specific deployment environment. # Operations Guides # Adopting an Existing Kubernetes Cluster as a Deployment Cell on Omnistrate Note If you are onboarding a customer-owned Kubernetes cluster for a fully managed deployment experience, use the [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) flow. With this, you can onboard your customer's Kubernetes cluster through an installation kit that deploys an agent to enable remote management of your deployments in that cluster. ## What Does It Mean to Adopt a Deployment Cell? Adopting a deployment cell means integrating an existing Kubernetes cluster into Omnistrate's cellular architecture as a managed deployment target. This process transforms your existing infrastructure into an Omnistrate-managed deployment cell while preserving your current investments and operational practices. ### Key Benefits of Adoption **Leverage Existing Infrastructure**: Instead of provisioning new infrastructure, you can utilize existing Kubernetes clusters that you or your customers have already invested in. This approach maximizes ROI on current infrastructure while gaining Omnistrate's management capabilities. **Maintain Control and Compliance**: Adopted cells allow you to keep infrastructure within your existing security boundaries and compliance frameworks. This is particularly valuable for: - Regulated industries with strict data residency requirements - Organizations with existing security and compliance certifications - Customers who prefer to maintain direct control over their infrastructure ### How Adoption Works When you adopt a deployment cell, Omnistrate installs a lightweight agent on your Kubernetes cluster that: - Establishes secure communication with the Omnistrate control plane - Enables deployment and management of services through Omnistrate - Provides monitoring and observability integration - Maintains the cluster's existing security and network configurations The adopted cluster becomes a fully managed deployment cell within Omnistrate's cellular architecture, allowing you to deploy and manage services alongside other deployment cells in your fleet. ## Deployment Cell Adoption Guide This guide walks you through the process of adopting an existing Kubernetes cluster as a deployment cell on Omnistrate using the `omctl` command-line tool. This allows you to onboard existing deployments in your or your customer's Kubernetes clusters and map them to deployment instances in the Omnistrate platform. ## Prerequisites - `omctl` CLI tool [installed](https://ctl.omnistrate.cloud/install/) - Access to the Kubernetes cluster you want to adopt ## Step-by-Step Guide ### 1. Login as SaaS Provider First, authenticate as the SaaS Provider who will be onboarding customer Kubernetes clusters: ``` omctl login ``` ### 2. Adopt the Kubernetes Cluster Use the `deployment-cell adopt` command to register your existing Kubernetes cluster: ``` omctl deployment-cell adopt \ --cloud-provider aws \ --description "adopted cluster" \ --id "cluster-1" \ --region us-east-1 \ --customer-email alok+drprod@omnistrate.com ``` **Parameters:** - `--cloud-provider`: The cloud provider where your cluster is hosted (e.g., `aws`, `gcp`, `azure`) - `--description`: A descriptive name for your deployment cell - `--id`: A unique identifier for the deployment cell - `--region`: The region where the cluster is located - `--customer-email`: (Optional) Email of the customer who owns the cluster. If not provided, the cluster will be adopted under the logged-in user This command will initiate the adoption process and register the cluster with Omnistrate. It will also download an installation kit for the agent that needs to be installed on the cluster in the same directory where you run the command. **Example output:** ``` Adopting deployment cell: ID: cluster-1 Cloud Provider: aws Region: us-east-1 Description: adopted cluster User Email: alok+drprod@omnistrate.com Deployment cell adoption initiated successfully! Adoption result: Adoption Status: PENDING_ADOPTION Agent Installation Kit: Saved as cluster-1.tar Note: Use the agent installation kit to complete the adoption process ``` ### 3. Install the Agent on the Cluster To complete the adoption, you need to install the Omnistrate agent on your Kubernetes cluster by following the README instructions provided in the downloaded installation kit (`cluster-1.tar`). Extract the kit and follow the instructions: ``` mkdir kit && cp cluster-1.tar kit && cd kit tar -xf cluster-1.tar # Eg: # First set the kubectl context to your desired cluster # Then apply the necessary Kubernetes manifests to set up the deployment agent: # Create namespace first and wait for it to be created kubectl apply -f ./ns.yaml && \ kubectl wait --for=jsonpath='{.status.phase}'=Active namespace/dataplane-agent --timeout=60s # Apply the necessary manifests to set up the deployment agent kubectl apply -f ./cluster-role-binding.yaml \ -f ./deployment.yaml \ -f ./dp-agent-tls.yaml \ -f ./priority-class.yaml \ -f ./sa.yaml # Wait for the agent to be ready kubectl -n dataplane-agent wait deployment/dp-agent --for=condition=Available --timeout=60s ``` ### 3. Check Adoption Status Verify the adoption status of your cluster: ``` omctl deployment-cell status --id cluster-1 ``` **Example output:** ``` +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ | CLOUD_PROVIDER | CURRENT_NUMBER_OF_DEPLOYMENTS | CUSTOMER_EMAIL | CUSTOMER_ORGANIZATION_NAME | HEALTH_STATUS | ID | REGION | STATUS | TYPE | +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ | aws | 0 | alok+drprod@omnistrate.com | Omnistrate | Status: UNKNOWN | Entities: 0/0 healthy | cluster-1 | us-east-1 | PENDING_ADOPTION | Kubernetes | +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ ``` ### 4. Filter Status by Customer Email To view the status for a specific customer: ``` omctl deployment-cell status --id cluster-1 --customer-email alok+drprod@omnistrate.com ``` **Example output:** ``` +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ | CLOUD_PROVIDER | CURRENT_NUMBER_OF_DEPLOYMENTS | CUSTOMER_EMAIL | CUSTOMER_ORGANIZATION_NAME | HEALTH_STATUS | ID | REGION | STATUS | TYPE | +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ | aws | 0 | alok+drprod@omnistrate.com | Omnistrate | Status: UNKNOWN | Entities: 0/0 healthy | cluster-1 | us-east-1 | PENDING_ADOPTION | Kubernetes | +----------------+-------------------------------+----------------------------+----------------------------+-----------------------------------------+-----------+-----------+------------------+------------+ ``` ### 5. Delete a Deployment Cell To remove a deployment cell: #### Deregistering a Deployment Cell ``` omctl deployment-cell delete --id cluster-1 --customer-email alok+drprod@omnistrate.com ``` You'll be prompted for confirmation: ``` Are you sure you want to delete deployment cell 'cluster-1'? This action cannot be undone. Type 'yes' to confirm: ``` For non-interactive deletion, use the `--force` flag: ``` omctl deployment-cell delete --id cluster-1 --customer-email alok+drprod@omnistrate.com --force ``` **Success output:** ``` Deleting deployment cell: cluster-1 Deployment cell 'cluster-1' deleted successfully! ``` #### Uninstalling the Agent After deregistering, you may want to uninstall the agent from the Kubernetes cluster. Follow the instructions in the README of the installation kit to remove the agent resources from your cluster. ## Important Notes - The `PENDING_ADOPTION` status indicates that the cluster adoption process is in progress and waiting for the agent to be installed and connected. - Multiple customers can have deployment cells with the same ID in different organizations - Deletion is permanent and cannot be undone - Health status will show as `UNKNOWN` until the adoption process is complete and the cluster is fully integrated ## Next Steps After successful adoption, you can: - Deploy services to the adopted cluster - Monitor cluster health and performance - Manage customer access and permissions For more information, refer to the Omnistrate documentation or contact support. # Managing and Upgrading Existing Customer Deployments with Omnistrate Omnistrate provides a powerful feature to adopt, manage, and upgrade existing customer application installations directly from the Omnistrate Command Line Interface (CLI). This allows your support and operations teams to centralize the management of both new and existing deployments, streamlining the entire lifecycle from a single control plane. This document outlines the process of adopting an existing customer deployment (specifically a Helm-based application running on a Kubernetes cluster) and managing its upgrade path. Note Currently, Omnistrate supports adopting Helm-based applications running on Kubernetes clusters. Future versions will expand support to other deployment types. ______________________________________________________________________ ## Core Concepts - **Deployment Cell**: In Omnistrate, a Deployment Cell is a fundamental building block representing an isolated environment where your application runs. In this context, a customer's existing Kubernetes cluster is treated as a Deployment Cell. - **Adoption**: This is the process of bringing an existing, unmanaged customer cluster and application under Omnistrate's management. This is a two-step process: first adopting the cluster (Cell), and then adopting the application (Instance) running within it. - **Omnistrate Agent**: A lightweight agent that gets installed in the customer's cluster. It establishes a secure, reverse-tunnel connection back to your control plane, allowing for management without requiring the customer to open inbound firewall ports. - **Instance**: Represents a specific application deployment, such as a Helm release, that has been adopted and is now managed by Omnistrate. ______________________________________________________________________ ## Step-by-Step Guide ### Step 1: Create a Customer User (if not already done) Before adopting a customer's deployment, ensure that the customer has a registered user account in the Omnistrate UI. This is necessary to link the deployment cell to their account. If the customer does not have an account, you can create one using the Omnistrate UI. Note Make sure to enable `Auto Verify User` to ensure the customer account is activated immediately. ### Step 2: Adopt the Customer's Kubernetes Cluster (Deployment Cell) The first step is to register the customer's existing Kubernetes cluster with Omnistrate. This is done using the `omctl deployment-cell adopt` command. For more details see the [Adopt Deployment Cells Guide](https://docs.omnistrate.com/operate-guides/adopt-deployment-cells/index.md). 1. **Run the Adopt Command**: Execute the following command, providing details about the customer's cluster. ``` omctl deployment-cell adopt \ --cloud-provider aws \ --description "Adopted Customer Cluster" \ --id "example-cluster" \ --region "us-east-1" \ --customer-email "customer.email@example.com" ``` - `--id`: A unique, memorable name for this deployment cell. - `--cloud-provider` & `--region`: The location of the customer's cluster. Can be omitted for on-premises clusters. - `--customer-email`: The email of the customer as registered in the Omnistrate portal, linking the cell to their account. 1. **Provide the Agent Installation Kit to the Customer**: Running the command successfully initiates the adoption process and generates an `Agent Installation Kit` (e.g., `example-cluster.tar`). This compressed file contains the necessary manifests for the customer to install the Omnistrate agent. The status of the cell will be `PENDING_ADOPTION` until the agent is installed. ### Step 3: Customer Installs the Omnistrate Agent The customer must perform this one-time installation in their cluster. 1. The customer untars the installation kit (`example-cluster.tar`). 1. The kit includes a `README.md` file with simple `kubectl` commands. 1. The customer runs the provided `kubectl apply` commands, which creates a new namespace (e.g., `dataplane-agent`) and deploys the Omnistrate agent. ### Step 4: Verify Cluster Adoption Once the agent is running, it will connect back to your Control Plane. You can verify this from your CLI: ``` omctl deployment-cell status --id example-cluster ``` The `STATUS` should now show as `HEALTHY` and `RUNNING`, indicating that Omnistrate can now manage the cluster. ### Step 5: Adopt the Application Instance (Helm Chart) With the cluster managed, you can now adopt the specific application running inside it. 1. **Create an Adoption Configuration File**: Create a YAML file (e.g., `adoption-config.yaml`) to describe the Helm release(s) you want to adopt. Note This file should include the Helm chart repository, release name, namespace, and any necessary credentials. Please ensure that the credentials are set correctly to allow the deployment cell to access the Helm chart repository. This example includes multiple Helm charts for the application platform version. ``` resourceAdoptionConfiguration: mainApp: helmAdoptionConfiguration: chartRepoURL: "oci://registry-1.docker.io/mycompanydev" releaseName: "app" releaseNamespace: "my-app" username: "admin" password: "dckr_pat_xyz123" runtimeConfiguration: wait: false waitForJobs: false supportService: helmAdoptionConfiguration: chartRepoURL: "oci://registry-1.docker.io/mycompanydev" releaseName: "support-svc" releaseNamespace: "my-app" username: "admin" password: "dckr_pat_xyz123" runtimeConfiguration: wait: false waitForJobs: false ``` 1. Get a list of all the supported platform versions for this service on Omnistrate ``` omctl service-plan list-versions --service-id s-xK8FMLq7R9 --plan-id pt-nB3Wz6yhKp -f="environment:Prod,service_name:My Application Platform" ✓ Service plan versions retrieved successfully +-------------+---------------+-----------------------------------+---------------------+--------------+-------------------------+---------+--------------------+ | ENVIRONMENT | PLAN_ID | PLAN_NAME | RELEASE_DESCRIPTION | SERVICE_ID | SERVICE_NAME | VERSION | VERSION_SET_STATUS | +-------------+---------------+-----------------------------------+---------------------+--------------+-------------------------+---------+--------------------+ | Prod | pt-nB3Wz6yhKp | My Application Enterprise | 3.2.1 | s-xK8FMLq7R9 | My Application Platform | 3.0 | Preferred | | Prod | pt-nB3Wz6yhKp | My Application Enterprise | 3.1.4 | s-xK8FMLq7R9 | My Application Platform | 2.0 | Active | | Prod | pt-nB3Wz6yhKp | My Application Enterprise | 2.5.8 | s-xK8FMLq7R9 | My Application Platform | 1.0 | Active | +-------------+---------------+-----------------------------------+---------------------+--------------+-------------------------+---------+--------------------+ ``` 1. **Run the Instance Adopt Command**: Execute the `instance adopt` command, referencing the cluster and the configuration file. The following command adopts a 2.5.8 application installation running on the `example-cluster` deployment cell, using the `adoption-config.yaml` file to specify the Helm charts and their configurations. ``` omctl instance adopt \ --host-cluster-id example-cluster \ --service-id s-xK8FMLq7R9 \ --service-plan-id pt-nB3Wz6yhKp \ --service-plan-version 1.0 \ --primary-resource-key mainApp \ -f adoption-config.yaml ``` This command registers the specified Helm charts as a single manageable instance. Omnistrate automatically discovers and stores the currently deployed Helm values. #### Step 6: Manage and Upgrade the Adopted Instance Once adopted, the application instance is a native deployment instance on Omnistrate. - **Secure Remote Access**: To debug or run pre-upgrade scripts, you can securely access the customer's cluster from your local machine. For more details, see the [Remotely Access Deployment Cells Guide](https://docs.omnistrate.com/operate-guides/deployment-cell-access/index.md). ``` omctl deployment-cell update-kubeconfig example-cluster --role cluster-admin ``` This command fetches a temporary `kubeconfig` by establishing a secure reverse-tunnel. You can now use `kubectl` and `helm` as if you were directly on the cluster, without any network changes on the customer's side. - **Perform an Upgrade**: Upgrading the instance is now a standard Omnistrate procedure. - **Generate both existing chart values and the proposed chart values for the upgrade.** The following command generates both the existing chart values and the chart values proposed from the target tier version as setup in your service template on Omnistrate. ``` # MAKE SURE TO REPLACE THE INSTANCE ID WITH THE ACTUAL INSTANCE ID FROM THE PREVIOUS ADOPTION STEP omctl instance version-upgrade --existing-configuration existing.yaml --proposed-configuration proposed-config.yaml --generate-configuration --target-tier-version 1.0 # CHANGE THE TARGET TIER VERSION TO MATCH THE TARGET PLATFORM VERSION ``` - **Run any pre-upgrade scripts or checks as needed.** Using the Secure Remote Access feature, you have direct access to the Kubernetes cluster to run any necessary pre-upgrade scripts or checks. - **Upgrade the instance with the default template values.** You can trigger an upgrade to a new version with the default args and the platform will use the values setup in the service template for that particular platform version. ``` omctl instance version-upgrade \ --target-tier-version 2.0 # Eg: Upgrading from 2.5.8 (Tier version 1.0) to 3.1.4 (Tier version 2.0) ``` - **Upgrade the instance with modified values.** If you need to override the default values, you can provide a custom values file. This allows you to specify any modifications needed for the upgrade. For example, if you have a modified values file named `modified-values.yaml`, you can run the following command: ``` omctl instance version-upgrade \ --target-tier-version 2.0 \ --upgrade-configuration-override modified-values.yaml ``` - **Applying modifications to existing values for the same platform version.** If you need to modify the existing values for the same platform version, you can use the same command as above but specify the target tier version same as the current version. ``` omctl instance version-upgrade \ --target-tier-version 1.0 \ --upgrade-configuration-override modified-values.yaml ``` - **Debug any installation issues.** If the upgrade fails or is taking too long, you can either use the Secure Remote Access to directly check the status using `kubectl` or `helm` commands. In addition, the helm client logs and the actual chart values being applied are available through the `instance debug` command: ``` omctl instance debug ``` - **Centralized View**: Use `omctl instance list` to see a consolidated list of all customer deployments, both newly created and adopted, providing a single pane of glass for managing your entire fleet. # Alarms ## Overview Alarm configuration is a global setting that allows an organization to subscribe to various Alerts and/or Notifications and forward them to supported notification channels. Notification settings can be modified in the UI under "Manage Account" -> "Alarm Settings". Each channel offers four types of filters: - Environment type - allows filtering based on environment type. Example: PROD / QA / DEV - Alert Event Type - allows filtering based on type of alert. Example: Alert / Notification - Alert Event Priority - allows filtering based on priority of alert. Example: Low / Medium / High / Critical - Event Category - allows filtering based on category. Example: InstanceEvents / UserEvents / DeploymentCellEvents, etc. - Event Type - allows filtering based on specific event type. Example: UnhealthyInstance / UserSignUp / DeploymentCellDeleteStarted, etc. - Payload - a json structure with properties of the event. Example: instance_id, subscription_id Each organization can configure multiple channels with different event filters. ## Interpreting instance health alerts Alerts complement, but do not replace, the health state shown on an instance. In practice: - `UnhealthyInstance` means one or more required health signals failed. - A running pod does not automatically mean the instance is healthy. - Endpoint reachability, readiness, custom health checks, and dependent integrations can all affect alerting and health state. When you receive an instance health alert, check: 1. The instance health details in the Operations Center 1. The workflow that last changed the instance 1. Instance debug output for rendered Helm or Terraform artifacts For more detail on health-state semantics, see [Monitoring](https://docs.omnistrate.com/operate-guides/monitoring/index.md) and [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md). ## Supported notification channels ### Email Email channel sends notification as email message to email address provided in email channel configuration. ### PagerDuty PagerDuty channels sends notification as PagerDuty alert. Recipient is identified by PagerDuty integration key that needs to be provided during configuration. For more details on how to generate integration key, refer to [this PagerDuty documentation](https://support.pagerduty.com/main/docs/services-and-integrations#generate-a-new-integration-key). ### Webhook Webhook calls HTTP-based callback function with configurable payload. By default, Omnistrate will include following details as request body: ``` { "eventID": "{{ $var.id }}", "serviceID": "{{ $var.ServiceID }}", "eventName": "{{ $var.Name }}", "eventDescription": "{{ $var.Description }}", "eventType": "{{ $var.Type }}", "payload": "{{ $var.Payload }}" } ``` Webhook channel allows method, endpoint and payload to be provided. If webhook action fails, we will retry few times before failing to notify this channel of the event. #### Signing webhook deliveries An endpoint that accepts any POST it receives will act on anything that reaches it. Give the channel a **signing secret** and every delivery arrives with an HMAC-SHA256 signature your receiver can check, so it can tell a real alarm from a request that merely knows the URL. Set it when you create or edit a Webhook channel: The secret must be at least **32 bytes**. That is not an arbitrary minimum: deliveries are signed with HMAC-SHA256, and [RFC 2104](https://www.rfc-editor.org/rfc/rfc2104#section-3) discourages a key shorter than the hash output because it weakens the function. A short key still produces a perfectly valid signature — it is simply cheaper to forge than the algorithm's name suggests. It is stored encrypted and **never returned by any read**, so keep your copy. A configured channel shows that signing is on, and an identifier for the active secret, so you can confirm a rotation took effect without the secret itself ever being displayed. ##### Headers on every delivery Five headers accompany a signed delivery, alongside `Content-Type: application/json` and any static headers you configured: | Header | Example | What it is for | | ------------------------------- | ---------------------- | ------------------------------------------------------------- | | `X-Omnistrate-Signature-256` | `sha256=24be8a3e…` | HMAC-SHA256 of the signed payload, lowercase hex | | `X-Omnistrate-Event-Id` | `evt-3c3d521bd975` | The event's own id. Stable across retries — deduplicate on it | | `X-Omnistrate-Event-Type` | `FailedBackup` | The alarm event type, matching the table below | | `X-Omnistrate-Timestamp` | `2026-08-25T07:58:45Z` | RFC 3339 UTC, taken when the request is sent | | `X-Omnistrate-Delivery-Attempt` | `1` | Starts at 1 | A complete request looks like this: ``` POST /omnistrate/alarms HTTP/1.1 Host: hooks.example.com Content-Type: application/json X-Omnistrate-Event-Id: evt-3c3d521bd975 X-Omnistrate-Event-Type: FailedBackup X-Omnistrate-Timestamp: 2026-08-25T07:58:45Z X-Omnistrate-Delivery-Attempt: 1 X-Omnistrate-Signature-256: sha256=24be8a3ec41a191588bb4a28cf5cef833f8fcef9050eada0fbd2d2b140b292b8 {"eventID":"evt-3c3d521bd975","serviceID":"s-Xy2ZaBcDeF","eventName":"FailedBackup", ...} ``` Your own configured headers are applied **first**, so a header you set cannot displace the signature. ##### Verifying the signature Sign the timestamp together with the body The signed material is `` + `.` + `` — **not the body alone**. Binding the timestamp in is what lets you enforce a freshness window: without it, someone who captures one delivery can replay it forever under a fresh timestamp and the signature still verifies. Verify against the **raw bytes**, before any parsing or re-serialization. Pretty-printing or re-encoding the JSON changes the bytes and the digest will not match. ``` const timestamp = req.headers['x-omnistrate-timestamp']; const signature = req.headers['x-omnistrate-signature-256']; if (typeof timestamp !== 'string' || typeof signature !== 'string') { return res.status(401).end(); } // rawBody is the unparsed request body. In Express, use express.raw() or capture it in a verify hook. const signed = `${timestamp}.${rawBody}`; const expected = 'sha256=' + crypto .createHmac('sha256', process.env.OMNISTRATE_WEBHOOK_SECRET) .update(signed, 'utf8') .digest('hex'); // Constant time, so a mismatch cannot be found one character at a time. const expectedBuf = Buffer.from(expected, 'utf8'); const signatureBuf = Buffer.from(signature, 'utf8'); if (expectedBuf.length !== signatureBuf.length || !crypto.timingSafeEqual(expectedBuf, signatureBuf)) { return res.status(401).end(); } // Then reject anything too old to be a live alarm. Five minutes is a reasonable window. if (Date.now() - Date.parse(timestamp) > 5 * 60 * 1000) { return res.status(401).end(); } ``` ##### What your receiver should do - **Answer quickly with any 2xx**, and do the work asynchronously. A receiver that blocks holds the delivery open. - **Deduplicate on `X-Omnistrate-Event-Id`.** It is stable across retries and across an operator re-sending an event, which is deliberate: re-sending exercises your idempotency rather than bypassing it. - **Use HTTPS.** Redirects are not followed — following one on a signed POST would either strip the body and signature or replay them at a host you never registered. The same signing scheme is used for marketplace fulfillment webhooks. See [Marketplace fulfillment API](https://docs.omnistrate.com/marketplace-fulfillment/api/index.md) if you integrate with those as well. ## Channels configuration Omnistrate generates a variety of Notifications and Alerts whenever a corresponding event occurs. Each notification channel can be configured to receive only specific types (based on event type, environment, priority, etc.). To add new channel, open "Manage account" -> "Notifications" page on UI where you will be able to add new channel. There are 2 configuration options available when adding a channel: 1. **Basic** - subscription is created based on environment type, alert event type and event priority. This option offers less control, but makes it easier to create a channel based on fewer inputs. All of basic dimensions (environment, type and priority) have limited set of options that are unlikely to change as we add new types of alerts. An example of a basic notification rule: "High and critical priority Alerts from prod environments". 1. **Advanced** - subscription is created based on environment type and specific category and/or event type. This option offers more control, but is based on event categories. ## Alarm Event Categories and Type The following table lists all available alarm event categories and the specific event types within each. For the payload delivered with each event, see [Webhook Alarm Event Payloads](https://docs.omnistrate.com/operate-guides/webhook-payloads/index.md). | Category | Type | Description | | -------------------------- | ---------------------------------- | --------------------------------------------- | | **InstanceEvents** | FailedBackup | Instance backup operation failed | | | FailedDelete | Instance deletion failed | | | FailedDeployment | Instance deployment failed | | | StartedDeployment | Instance deployment process initiated | | | FailedRecovery | Instance recovery operation failed | | | FailedRestore | Instance restore operation failed | | | FailedRestart | Instance restart failed | | | FailedSnapshotCopy | Instance snapshot copy operation failed | | | FailedSnapshotCreate | Instance snapshot creation failed | | | FailedSnapshotDelete | Instance snapshot deletion failed | | | FailedStart | Instance start operation failed | | | FailedStop | Instance stop operation failed | | | FailedUpdate | Instance update operation failed | | | HighCPUUsage | Instance CPU usage exceeds threshold | | | RecoveryStarted | Instance recovery process initiated | | | ScaleDownFailed | Instance scale down operation failed | | | ScaleDownSuccess | Instance scale down completed successfully | | | ScaleUpFailed | Instance scale up operation failed | | | ScaleUpSuccess | Instance scale up completed successfully | | | StartedDelete | Instance deletion process initiated | | | StartedRestore | Instance restore process initiated | | | SuccessfulBackup | Instance backup completed successfully | | | SuccessfulDelete | Instance deleted successfully | | | SuccessfulDeployment | Instance deployed successfully | | | SuccessfulRecovery | Instance recovery completed successfully | | | SuccessfulRestore | Instance restore completed successfully | | | SuccessfulRestart | Instance restarted successfully | | | SuccessfulSnapshotCopy | Instance snapshot copy completed successfully | | | SuccessfulSnapshotCreate | Instance snapshot created successfully | | | SuccessfulSnapshotDelete | Instance snapshot deleted successfully | | | SuccessfulStart | Instance started successfully | | | SuccessfulStop | Instance stopped successfully | | | SuccessfulUpdate | Instance updated successfully | | | UnhealthyCustomerIntegration | Customer integration is unhealthy | | | UnhealthyInstance | Instance health check failed | | | UnhealthyIntegration | Integration is unhealthy | | **UserEvents** | ApproveSubscriptionRequest | Subscription request approved | | | UserSignUp | New user registration | | | UserSubscription | User subscription created | | | UserSubscriptionInvite | User invited to subscription | | | UserSubscriptionRevoked | User subscription access revoked | | | UserUnsubscribed | User unsubscribed from service | | | UserDeleted | User account deleted | | **IdentityProviderEvents** | FailedIdentityProviderVerification | Identity provider verification failed | | **SystemEvents** | UpgradeScheduled | System upgrade scheduled | | | UpgradeMaintenanceActionRequest | Maintenance action requested for upgrade | | | UpgradePaused | System upgrade paused | | **BillingEvents** | S3MeteringExportFailed | S3 metering export operation failed | | | GCSMeteringExportFailed | GCS metering export operation failed | | | InvoiceGenerateSuccess | Invoice generation completed successfully | | **DeploymentCellEvents** | DeploymentCellCreateCompleted | Deployment cell creation finished | | | DeploymentCellCreateStarted | Deployment cell creation initiated | | | DeploymentCellDeleteCompleted | Deployment cell deletion finished | | | DeploymentCellDeleteStarted | Deployment cell deletion initiated | | | DeploymentCellUpdateCompleted | Deployment cell update finished | | | DeploymentCellUpdateStarted | Deployment cell update initiated | | | DeploymentCellCompleted | Deployment cell operation completed | | | DeploymentCellStarted | Deployment cell operation started | | | DeploymentCellInProgress | Deployment cell operation in progress | | | HostClusterCleanup | Host cluster cleanup operation | | | RepairingDeploymentCellStarted | Deployment cell repair process initiated | # AWS CloudFormation Account Controls Use this guide when the customer account owner adjusts CloudFormation controls in the customer-owned CloudFormation onboarding stack, `AccountConfigSetup`, to set: - `K8sDebugAccessEnabled`, which controls whether dataplane agent pods in BYOC PrivateLink accounts can reach the Manager Kubernetes API proxy port used for Kubernetes debug access to the customer Kubernetes API. - `AgentInfrastructureMutationEnabled`, which controls whether the dataplane agent can make AWS infrastructure changes through the account-config roles. The update runs in the AWS account that owns the stack and is visible in AWS CloudFormation history. ## Parameters | Parameter | Value | Result | | ------------------------------------ | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `K8sDebugAccessEnabled` | `true` | BYOC PrivateLink accounts only. Removes the `block-k8s-api-proxy` `NetworkPolicy` from the `dataplane-agent` namespace, allowing `dataplane-agent` egress to the manager Kubernetes API proxy port. | | `K8sDebugAccessEnabled` | `false` | BYOC PrivateLink accounts only. Applies the `block-k8s-api-proxy` `NetworkPolicy` in the `dataplane-agent` namespace to block `dataplane-agent` egress to the manager Kubernetes API proxy port. | | `AgentInfrastructureMutationEnabled` | `true` | Allows the dataplane agent to create, update, and delete AWS infrastructure when required for provisioning, updates, deletes, and reconciliation. | | `AgentInfrastructureMutationEnabled` | `false` | Disables dataplane agent AWS infrastructure mutation by applying an explicit deny for non-read AWS actions while preserving read-only inspection. | `K8sDebugAccessRegions` is optional and only applies to BYOC PrivateLink debug-access reconciliation. Use it only when you need to limit reconciliation to specific AWS regions or `region:host-cluster-id` targets, for example `us-west-2:hc-xxxx,us-east-1`. Leave it blank to scan all enabled regions. ## Update the Stack 1. Open the generated AWS CloudFormation update URL or command for the target AWS account. 1. Keep generated parameters unchanged, including OIDC values, service account values, `ProvisionAccountConfig=true`, and generated account-type or topology values such as `IsBYOAAccount` and `IsBYOCPrivateAccount`. 1. Change `K8sDebugAccessEnabled`, `AgentInfrastructureMutationEnabled`, or both. 1. Submit the stack update and wait for `UPDATE_COMPLETE`. ## Verify the State After CloudFormation reaches `UPDATE_COMPLETE`, verify the stack parameters in the AWS CloudFormation console. The Kubernetes debug-access reconciliation runs asynchronously after the stack update is queued. Allow that reconciliation to finish before checking the Kubernetes `NetworkPolicy`; `UPDATE_COMPLETE` does not by itself guarantee that the policy has already been applied or removed. For `K8sDebugAccessEnabled`, the blocked state corresponds to the `block-k8s-api-proxy` `NetworkPolicy` being present in the `dataplane-agent` namespace. The unblocked state corresponds to that policy being absent. For dataplane agent infrastructure mutation, verify that `AgentInfrastructureMutationEnabled` matches the intended value. # Manage BYOC Cloud Accounts View and manage your customers' BYOC Cloud Accounts for BYOC deployments. ## What is BYOC? BYOC (Bring Your Own Cloud) is a deployment model where your customers' applications are deployed and managed within their own cloud accounts rather than in your infrastructure. This approach addresses critical customer requirements around data sovereignty, security compliance, and cost control while still providing a fully-managed SaaS experience. ## BYOC Operations ### Self-served BYOC Setup The **[Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md)** can be used by customers to connect and manage their BYOC cloud accounts. ### Setup Customer Accounts on behalf of customer You can also work with your customer to perform an assisted setup of their Cloud Account. ## Responsibility model In a BYOC deployment, responsibilities are split across three parties: - **Customer account owner**: Runs the onboarding flow in the target account, approves IAM and identity setup, and owns cloud quotas, organization policies, and network constraints in that account. - **SaaS provider**: Enables BYOC for the service, guides or assists the customer through onboarding, and operates the resulting instances through Omnistrate. - **Omnistrate**: Provides the onboarding artifacts, bootstraps the deployment cell, installs platform components, and orchestrates lifecycle operations in the connected account. ## Lifecycle of a BYOC account For most BYOC services, the lifecycle looks like this: 1. The customer connects their cloud account from the Customer Portal or through an assisted flow. 1. The first instance in a given account and region bootstraps the deployment cell for that location. 1. Later instances in the same account and region typically reuse that deployment cell instead of creating a new cluster every time. 1. The account can only be offboarded after all instances in that account are deleted. For more background on deployment cells, see [Deployment Cells](https://docs.omnistrate.com/build-guides/deployment-cells/index.md). ## Self-serve vs assisted onboarding Both patterns are supported: - **Self-serve**: The customer runs the onboarding flow directly from the Customer Portal. - **Assisted**: Your team coordinates with the customer and helps execute the onboarding steps in their cloud account. The underlying Omnistrate bootstrap flow is the same in both cases. The main difference is who drives the cloud-console or Cloud Shell steps. ## CTL workflow for BYOA cloud accounts If you prefer to onboard customer accounts from the CLI, use the provider-specific getting-started guides: - [AWS account onboarding with CTL](https://docs.omnistrate.com/getting-started/onboarding/aws/index.md) - [GCP account onboarding with CTL](https://docs.omnistrate.com/getting-started/onboarding/gcp/index.md) - [Azure account onboarding with CTL](https://docs.omnistrate.com/getting-started/onboarding/azure/index.md) - [Nebius account onboarding with CTL](https://docs.omnistrate.com/getting-started/onboarding/nebius/index.md) - [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) for customer-owned Kubernetes clusters The common BYOA lifecycle commands are: ``` # Create a customer onboarding instance omnistrate-ctl account customer create [provider flags] --service --environment --plan # List and inspect customer onboarding instances omnistrate-ctl account customer list [flags] omnistrate-ctl account customer describe # Update or delete a customer onboarding instance omnistrate-ctl account customer update [flags] omnistrate-ctl account customer delete # Create a BYOA deployment with that onboarding instance omnistrate-ctl instance create ... --customer-account-id ``` In production environments, `account customer create` defaults to the calling user's subscription unless you explicitly pass `--subscription-id` or `--customer-email`. ## Account Tags Cloud accounts support custom tags — free-form key/value pairs stored on the account configuration. The account configuration is the definitive store for these tags: every deployment cell running in the account exposes them as the [`$sys.deploymentCell.accountTags`](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/#conditioning-on-cloud-account-tags) system parameter, which can steer [conditional amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/#conditional-amenities), amenity Helm chart values, and service plan configuration. ### Setting tags with CTL All account command sets accept a `--tags` flag with comma-separated `key=value` pairs: ``` # Tag your own cloud account at onboarding time omnistrate-ctl account create my-account --aws-account-id=123456789012 \ --tags "env=prod,team=platform" # Replace the tag set on an existing account omnistrate-ctl account update my-account --tags "env=prod,omnistrate.com/cost-center=cc-1234" # Tag a customer BYOA account at onboarding time — the tags flow through the # onboarding instance to the customer's backing account configuration omnistrate-ctl account customer create --service --environment \ --plan --aws-account-id= --tags "tier=enterprise" # Replace the tag set on a customer's backing account configuration omnistrate-ctl account customer update --tags "tier=enterprise,region-pref=eu" ``` Tags appear in `account describe`, `account list`, and `account customer describe` output, and in the corresponding API responses (`customTags`). ### Semantics - **Update replaces the whole tag set.** `--tags` on `account update` and `account customer update` is a full replacement, not a merge — include every tag you want to keep. - **BYOA instance tags write to the account.** For BYOA account onboarding instances, custom tags set through the instance APIs (create or metadata update) are written to the backing account configuration, and reads return the account configuration's tags — the instance never keeps a separate tag set. An account shared by multiple onboarding instances has one tag set; the last write wins. - **Key format.** Tag keys may contain dots and slashes (for example `omnistrate.com/cost-center`); reference such keys with bracket notation in system parameter expressions. Keys used to filter [layered chart values scopes](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/#scoped-layer) must use only letters, digits, underscores, and dots. - **Duplicate keys are rejected**, both by the CLI and the API. ## BYOC On-Premise Kubernetes Accounts BYOC On-Premise connects a customer-owned Kubernetes cluster instead of a cloud account. For the full setup flow, see [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md). ## Retrying onboarding safely When onboarding fails partway through, or when bootstrap resources are deleted manually later, avoid trying to stitch old and new bootstrap state together. Recommended approach: - Re-run the current onboarding flow for the exact target account and cloud provider. - If the connected account was partially onboarded and then manually modified, disconnect or offboard it and onboard again cleanly. - For GCP, treat deleted-and-recreated projects, removed workload identity pools/providers, or mismatched onboarding kits as fresh onboarding events. - Do not assume a previous bootstrap in another region, another environment, or another cloud account can be reused automatically. ## CloudFormation Controls For AWS BYOC accounts, the customer account owner can use the customer-owned CloudFormation onboarding stack to manage account controls. `K8sDebugAccessEnabled` applies to BYOC PrivateLink support/debug access to the customer Kubernetes API, while `AgentInfrastructureMutationEnabled` controls AWS infrastructure mutation permissions on account-config roles. See [AWS CloudFormation Account Controls](https://docs.omnistrate.com/operate-guides/aws-cloudformation-account-controls/index.md). ## BYOC PrivateLink Accounts [BYOC PrivateLink](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md) is selected at account-onboarding time (Customer Portal toggle, or `--private-link` on `omnistrate-ctl account customer create`). Once enabled, every instance in the account uses the PrivateLink dataplane topology. ### Omnistrate-managed VPC The simplest topology: Omnistrate provisions the VPC, subnets, NAT gateway, security group, and the management VPC Endpoint inside the customer's account on first deployment. The customer only grants the bootstrap role via CloudFormation — no manual VPC or VPCE setup is required. To allow this path, the account must be onboarded with the **Allow new cloud-native network creation** option enabled. This grants Omnistrate permission to create new VPCs in the account on demand; without it, every deployment must reuse an existing customer-supplied VPC. - **Customer Portal:** enable **Allow new cloud-native network creation** on the AWS account form. - **`omnistrate-ctl`:** pass `--allow-create-new-cloud-native-network` on [`account customer create`](https://ctl.omnistrate.cloud/omnistrate-ctl_account_customer/): ``` omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --aws-account-id= \ --private-link \ --allow-create-new-cloud-native-network ``` When this option is set, the first instance in each region triggers Omnistrate to create the deployment cell's VPC, subnets, NAT gateway, security group, and management VPCE. Later instances in the same region reuse the same network. ### Imported VPC requirements for BYOC PrivateLink Customers typically bring their own VPC when they need to comply with internal network governance — for example, reusing an existing CIDR plan that fits Transit Gateway / Direct Connect routing, attaching the dataplane to a pre-approved security group and firewall posture, sharing a NAT gateway or egress proxy with other workloads, or operating under a policy that prohibits the SaaS provider from creating new VPCs in their account. When the customer brings their own VPC for a PrivateLink account, the VPC must satisfy the following before any instance is deployed: | # | Requirement | Details | | --- | ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | **VPC & subnet tags** | Tag the VPC and the subnets where deployment instances should run with `omnistrate.com/managed-by` = `omnistrate`. The tag is both the IAM permission gate and the subnet selector for EKS nodegroup placement — tag **only** the subnets where you want deployment instances to be placed. | | 2 | **DNS** | Enable `enableDnsSupport` and `enableDnsHostnames` on the VPC. | | 3 | **Egress** | Deployment subnets need outbound internet via NAT Gateway, Transit Gateway, or VPN. Required for Helm binary downloads and `ghcr.io` chart pulls during dataplane bootstrap. | | 4 | **Management VPC Endpoint** | Create a single Interface VPC Endpoint targeting the PrivateLink service Omnistrate provides, with: | | | | • Tag `Name` = `omnistrate-byoc-private-vpce-` | | | | • Security group allowing inbound TCP **8443–8506** from the VPC CIDR | | | | • Tag every security group attached to the VPCE with `omnistrate.com/managed-by` = `omnistrate`. Omnistrate uses this tag as the IAM permission gate when adding or removing Kubernetes debug access ingress rules. | | 5 | **Cross-region** *(if applicable)* | If the customer's VPC and the PrivateLink service live in different AWS regions, pass `--service-region` to `aws ec2 create-vpc-endpoint` (or `service_region` in the Terraform `aws_vpc_endpoint` resource). Do **not** enable private DNS — AWS does not support it for cross-region interface endpoints. | | 6 | **Subnet tags** *(optional)* | Private subnets: tag `kubernetes.io/role/internal-elb` = `1`. Public subnets: tag `kubernetes.io/role/elb` = `1`. Not mandatory for PrivateLink deployments, but recommended if you plan to use internal load balancers. | Note Omnistrate hands the customer the VPCE service name and the provisioner host-cluster ID after the account is onboarded; both are needed to create the management VPC Endpoint. ## Account Offboarding When a customer no longer needs your SaaS Product or wants to disconnect their cloud account, you can offboard their BYOC account. This removes all Omnistrate-managed resources from the customer's cloud account and severs the trust relationship. ### Prerequisites Before a BYOC account can be offboarded: 1. **The customer must delete all deployment instances** running in their cloud account. Active instances must be terminated before the account can be offboarded. Warning Offboarding is a destructive operation. All deployment cells, networking, and control plane components in the customer's cloud account are permanently removed. Ensure the customer has backed up any data they need to retain. ## Offboarding Steps ### Offboarding from the Customer Portal Customers offboard their BYOC account directly from the [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md): 1. Navigate to your cloud account settings in the Customer Portal. 1. Delete the cloud account configuration. 1. Confirm the offboarding operation. 1. Wait for Omnistrate to automatically remove all managed resources from your account. ### What offboarding cleans up When you offboard a customer's BYOC account, Omnistrate removes the following resources from the customer's cloud: - **Control plane agents** and system nodes - **Monitoring infrastructure** - **Networking resources** (VPCs, subnets, and security groups created by Omnistrate) — only if you are not using a Bring Your Own Network deployment model - **Deployment cells** and their underlying Kubernetes clusters ### Customer-side cleanup After the offboarding process completes, the customer should remove the onboarding permissions from their cloud account. For cloud-specific cleanup steps, see [Offboarding](https://docs.omnistrate.com/getting-started/account-onboarding/#offboarding). Note The cloud-provider-specific onboarding permissions are not automatically removed during offboarding. The customer must clean these up manually. # Remotely Access Any Deployment Cell (Kubernetes Cluster) with Omnistrate's Secure CLI Omnistrate is revolutionizing how service providers manage and interact with their deployment cells. A new feature that now provides a secure and standardized way to remotely access the Kubernetes API of any deployment cell (Kubernetes Cluster). This capability eliminates the need for complex networking configurations or additional permissions in the target account. This new functionality allows you to securely connect to any Kubernetes cluster, whether it's running in your own cloud account, a customer's account, or across different cloud providers and regions. ## Key Benefits - **Simplified Operations:** Gain direct access to any managed Kubernetes cluster with a single command. This removes the need for bastion hosts, VPNs, or managing separate credentials for each environment, dramatically simplifying day-to-day operations. - **Standardized Cross-Channel Management:** Whether you're debugging an issue in a development cluster or a production instance running in a customer's environment, the process is identical. This standardization reduces complexity and the risk of human error. - **Enhanced Security:** All connections are established over a secure mTLS (mutual TLS) channel using reverse tunneling. This means you don't need to expose the Kubernetes API server to the public internet or open inbound firewall ports in the target account. Access is further protected by short-term tokens, ensuring a robust Zero Trust security model. - **Seamless Tool Integration:** Your existing Kubernetes tooling, including `kubectl`, `k9s`, and `helm`, works out-of-the-box. Once the secure connection is established, you can use your preferred tools to manage and inspect your applications and services. ## How It Works: Secure by Design Omnistrate's remote access feature is built on a foundation of security and simplicity. When you initiate a remote session, the following happens: 1. **Secure Tunneling:** A secure reverse tunnel is created from your local machine to the target Kubernetes cluster. This connection is protected with mTLS, ensuring that both the client and server authenticate each other's identity through trusted certificates. 1. **Dynamic Configuration:** The Omnistrate CTL (`omnistrate-ctl`) dynamically updates your local `kubeconfig` file with temporary, short-lived credentials for the target cluster. 1. **Direct API Access:** With the `kubeconfig` updated, `kubectl` and other Kubernetes-native tools can securely communicate with the cluster's API server through the mTLS tunnel. This architecture ensures that access is granted on-demand and is automatically revoked, without ever exposing the cluster's control plane to external threats. ## Getting Started: A Step-by-Step Guide Interacting with your remote clusters is straightforward using the Omnistrate CLI. ### Step 1: List Your Deployment Cells To see all the deployment cells you manage, run the following command ([More details](https://ctl.omnistrate.cloud/omnistrate-ctl_deployment-cell_list/)): ``` omnistrate-ctl deployment-cell list +----------------+-------------------------------+---------------------+----------------------------+---------------------------------------------+--------------+-------------+---------+------------+ | CLOUD_PROVIDER | CURRENT_NUMBER_OF_DEPLOYMENTS | CUSTOMER_EMAIL | CUSTOMER_ORGANIZATION_NAME | HEALTH_STATUS | ID | REGION | STATUS | TYPE | +----------------+-------------------------------+---------------------+----------------------------+---------------------------------------------+--------------+-------------+---------+------------+ | azure | 1 | | | Status: UNKNOWN | Entities: 0/0 healthy | hc-x7n2kp9q4 | eastus2 | FAILED | Kubernetes | | gcp | 4 | | | Status: HEALTHY | Entities: 556/556 healthy | hc-3m8vt5hy2 | us-central1 | RUNNING | Kubernetes | | aws | 2 | | | Status: HEALTHY | Entities: 195/195 healthy | hc-p4j6wd3n7 | us-east-1 | RUNNING | Kubernetes | | aws | 1 | | | Status: HEALTHY | Entities: 139/139 healthy | hc-k9z1bf8c5 | us-west-2 | RUNNING | Kubernetes | | aws | 0 | | | Status: UNKNOWN | Entities: 0/0 healthy | hc-v2q7rx4m1 | ap-south-1 | FAILED | Kubernetes | | aws | 1 | demo@omnistrate.com | Omnistrate | Status: HEALTHY | Entities: 242/242 healthy | hc-h5y8sg2w6 | us-east-1 | RUNNING | Kubernetes | +----------------+-------------------------------+---------------------+----------------------------+---------------------------------------------+--------------+-------------+---------+------------+ ``` This command provides a comprehensive overview of your clusters, including their health status, region, and the customer account they belong to. ### Step 2: Connect to a Remote Cluster To establish a secure session with a specific deployment cell, use the `update-kubeconfig` command. You will need the ID of the deployment cell from the list in the previous step. If the cluster belongs to a customer, you will also need their email address. [More details](https://ctl.omnistrate.cloud/omnistrate-ctl_deployment-cell_update-kubeconfig/) ``` omnistrate-ctl deployment-cell update-kubeconfig hc-h5y8sg2w6 --customer-email demo@omnistrate.com ``` This command will securely update your local kubeconfig file. The path to the updated configuration will be displayed in your terminal. You can also specify a custom path for the kubeconfig file: ``` omnistrate-ctl deployment-cell update-kubeconfig hc-h5y8sg2w6 --customer-email demo@omnistrate.com --kubeconfig /path/to/custom/kubeconfig ``` ### Step 3: Interact with Your Cluster You are now securely connected to the remote Kubernetes cluster. You can use your favorite tools to manage your applications. For example, you can list the namespaces in the cluster: ``` # Set the KUBECONFIG environment variable to the path provided by the previous command export KUBECONFIG=/tmp/kubeconfig # Use kubectl to interact with the cluster # Check if you can create pods in the current namespace kubectl auth can-i create pods no # Check if you can create pods in a specific namespace kubectl auth can-i create pods --namespace=kube-system no ``` By default, the kubeconfig will assume a read-only cluster-wide role (`cluster-reader`). If you need to perform administrative tasks, you can specify a different role: ``` omnistrate-ctl deployment-cell update-kubeconfig hc-h5y8sg2w6 --customer-email demo@omnistrate.com --role cluster-admin ``` You can also use tools like **k9s** for a more interactive terminal UI to manage your cluster's resources, view logs, and much more, all through the secure Omnistrate tunnel. ## Deployment Cell Bootstrap Visibility When Omnistrate provisions a new deployment cell (Kubernetes cluster), you can monitor the bootstrap process and troubleshoot failures through the Operations Center. ### Debug events If a deployment cell fails to provision or encounters errors during bootstrap, Omnistrate surfaces debug events that help you identify the root cause: - **Bootstrap failures**: Errors during cluster creation (e.g., quota limits, permission issues) - **Add-on failures**: Issues with required Kubernetes add-ons or drivers - **Network configuration errors**: Problems with VPC or subnet setup - **Cloud provider errors**: API errors or resource limits from the underlying cloud provider You can view these events in the Operations Center workflow details for the specific deployment cell provisioning operation. Tip If a deployment cell bootstrap fails, check the workflow details for specific error messages. Common causes include cloud account quota limits, missing IAM permissions, or network CIDR conflicts. # Deployment Cell Amenities ## Overview Deployment cell amenities are additional infrastructure components (Helm charts) that can be installed and managed on your deployment cell to enhance functionality and provide essential services. Omnistrate provides both managed amenities and support for custom amenities to meet your specific infrastructure requirements. This guide demonstrates how to manage deployment cell amenities using the Omnistrate command-line tool. ## Prerequisites Before managing deployment cell amenities, ensure you have: - Omnistrate CTL installed ([installation guide](https://ctl.omnistrate.cloud/install/)) - Valid credentials for the Omnistrate platform - Appropriate permissions to manage deployment cell - Access to your target environment (e.g., PROD, STAGING) Deployment cell support two types of amenities: Amenities are the right place for components that are cluster-wide rather than instance-specific, such as monitoring stacks, CSI drivers, ingress controllers, shared Operators, and policy controllers. Managing them at the cell layer lets you upgrade them once per cluster instead of tying their lifecycle to every deployment instance. ## Managed Amenities Pre-configured Helm charts maintained by Omnistrate that are automatically available for installation. These amenities are: - **Fully managed**: Omnistrate handles chart configuration, updates, and maintenance - **Cloud-optimized**: Configured with best practices for each cloud provider - **Production-ready**: Tested and validated for enterprise use Common managed amenities include: - **Observability Prometheus**: Monitoring and metrics collection. This amenity installs the full Grafana observability stack on the deployment cell, including **Prometheus** for metrics, **Grafana** for dashboards and visualization, **Grafana Loki** for logs, and **Grafana Tempo** for traces. Logs and traces are collected and forwarded to Loki and Tempo respectively by Grafana Alloy collectors deployed alongside the stack. - **Kubernetes Dashboard**: Web-based Kubernetes cluster management interface - **External DNS**: Automatic DNS record management for Kubernetes services - **Cert Manager**: Automatic SSL certificate provisioning and renewal - **Nginx Ingress Controller**: HTTP/HTTPS traffic routing and load balancing Each cloud provider has additional managed amenities. ### Default Amenities Installed Per Cell When a deployment cell is bootstrapped, Omnistrate installs a baseline set of amenities so the cluster is production-ready out of the box. The table below shows which amenities are installed by default on each cloud provider (✅ = installed by default). | Amenity | AWS | GCP | Azure | OCI | | ------------------------------------------------------------------------------- | --- | --- | ----- | --- | | Observability Prometheus (full Grafana stack: Prometheus, Grafana, Loki, Tempo) | ✅ | ✅ | ✅ | ✅ | | Kubernetes Dashboard | ✅ | ✅ | ✅ | ✅ | | External DNS | ✅ | ✅ | ✅ | ✅ | | Cert Manager | ✅ | ✅ | ✅ | ✅ | | Nginx Ingress Controller | ✅ | ✅ | ✅ | ✅ | | Cluster Autoscaler | ✅ | ✅ | ✅ | ✅ | | EBS CSI Driver | ✅ | — | — | — | | EFS CSI Driver | ✅ | — | — | — | | S3 CSI Driver | ✅ | — | — | — | | AWS Load Balancer Controller | ✅ | — | — | — | | IP Masq Agent | — | ✅ | — | — | | KEDA (event-driven autoscaling) | — | ✅¹ | — | — | | Azure NVIDIA Device Plugin | — | — | ✅ | — | | CoreDNS addon | — | — | — | ✅ | | NVIDIA DCGM Exporter | ✅² | ✅² | ✅² | ✅² | ¹ Installed on GCP when event-driven autoscaling is enabled. ² Installed on any cloud when the cell is provisioned with NVIDIA GPU node pools; feeds GPU metrics into the observability stack. Note The default amenity set is managed by Omnistrate and configured with cloud-optimized defaults. You can extend or override the installed amenities for a given cloud provider through the deployment cell configuration template, as described in the following sections. ## Custom Amenities 1. Helm charts that you configure and manage yourself. 1. Plain Kubernetes Manifests that you can deploy and manage on your deployment cell. Each custom amenity is defined with these top-level fields: | Field | Type | Required | Description | | ------------- | ------ | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `name` | string | Yes | Unique name for the amenity within the template | | `description` | string | No | Human-readable description | | `type` | string | Yes | `Helm` or `KubernetesManifest` | | `properties` | object | Yes | Type-specific configuration. See [Helm Chart Amenities](#helm-chart-amenities) and [Kubernetes Manifest Amenities](#kubernetes-manifest-amenities) | | `dependsOn` | array | No | List of amenity names that must finish installing before this amenity starts. This is a top-level field — not inside `properties`. See [Amenity Dependency Ordering](#amenity-dependency-ordering) | | `disable` | string | No | Expression evaluated per deployment cell that skips the amenity when it renders to `true`. See [Conditional Amenities](#conditional-amenities) | Execution Model Omnistrate applies custom amenities in parallel by default. If one amenity depends on another resource being present first — such as a CRD, namespace, Secret, or operator — use the `dependsOn` field to declare the ordering. See [Amenity Dependency Ordering](#amenity-dependency-ordering) for details and examples. Alternatively, you can package prerequisite resources into the same Helm chart so the chart owns the ordering internally. Deletion and rollback use the same model The same rendering and parallelism rules apply when amenities are removed or when you revert to an older template. - `KubernetesManifest` amenities are rendered again during removal so Omnistrate can determine what to delete. - If a manifest contains unresolved expressions, unsupported function chains, or values that no longer exist, the delete or revert workflow can block the entire sync. - Keep manifest amenities deletion-safe. Prefer direct values or single secret substitutions rather than multi-step transforms. ## Managing Deployment Cell Configuration Templates Configuration templates define which amenities are available for a deployment cell on each cloud provider. Note By default Omnistrate uses a Configuration Template per cloud provider. If you need to define different templates per Environment please reach out to [support@omnistrate.com](mailto:support@omnistrate.com) Template Updates and Existing Cells Updating an organization-level amenities template changes the desired template for that cloud provider, but it does not immediately reconfigure existing deployment cells. To roll those changes out to an existing cell: 1. Sync the cell with the latest template using `omnistrate-ctl deployment-cell update-config-template --id --sync-with-template` 1. Apply the pending changes using `omnistrate-ctl deployment-cell apply-pending-changes --id ` ### Generate Configuration Template Create a template with all available amenities for a specific cloud provider: ``` omnistrate-ctl deployment-cell generate-config-template --cloud aws --output template-aws.yaml ``` When prompted, select your login method and provide credentials. ### Using Schema Validation To simplify the definition of this specification file Omnistrate provide a JSON schema that can be used for validation. You can use the JSON schema in IDEs that use the YAML Language Server (eg: VSCode / NeoVim). ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/deployment-cell-amenities-spec-schema.json ``` For [**IntelliJ**](https://www.jetbrains.com/help/idea/yaml.html#use-schema-keyword) replace the top line with the following line to set up the yaml schema ``` # $schema: https://api.omnistrate.cloud/2022-09-01-00/schema/deployment-cell-amenities-spec-schema.json ``` ### Review Template Structure The generated template contains two main sections: ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/deployment-cell-amenities-spec-schema.json managedAmenities: - name: EFS CSI Driver description: EFS CSI Driver type: Helm - name: AWS Load Balancer Controller description: AWS Load Balancer Controller type: Helm - name: Cluster Autoscaler description: Cluster Autoscaler type: Helm # ... additional managed amenities customAmenities: - name: EBS CSI Driver description: EBS CSI Driver type: Helm properties: ChartName: "aws-ebs-csi-driver" ChartVersion: "2.28.1" ChartRepoName: "aws-ebs-csi-driver" ChartRepoURL: "https://kubernetes-sigs.github.io/aws-ebs-csi-driver" CredentialsProvider: Type: "none" DefaultNamespace: "kube-system" ChartValues: image: repository: "public.ecr.aws/ebs-csi-driver/aws-ebs-csi-driver" tag: "v1.27.0" controller: replicaCount: 2 resources: requests: cpu: "10m" memory: "40Mi" affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/control-plane operator: Exists tolerations: - key: CriticalAddonsOnly value: "true" effect: NoSchedule ``` Warning Omnistrate creates a system node pool that is independent of the service plan node pool. To allow Operators and Controllers to run on the system node pool, affinity rules and tolerations must be defined as shown in the example and explained in the [Affinity rules and tolerations section](#affinity-rules-and-tolerations). ## Helm Chart Amenities When defining custom Helm amenities, the `properties` object supports: | Property | Description | Required | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | `ChartName` | Name of the Helm chart | Yes | | `ChartVersion` | Specific version of the chart | Yes | | `ChartRepoName` | Repository name identifier | Yes | | `ChartRepoURL` | Repository URL | Yes | | `CredentialsProvider` | Authentication configuration | No | | `DefaultNamespace` | Kubernetes namespace for deployment | No | | `ChartValues` | Custom Helm values to override defaults | No | | `LayeredChartValues` | Ordered list of value layers merged per deployment cell, each optionally gated by a `scope`. See [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) | No | | `ReleaseName` | Optional custom release name for Helm chart. This setting is useful in case you need to install multiple versions of the same chart | No | **Example: install the same chart twice** ``` customAmenities: - name: opentelemetry-collector-agents description: Node-level collectors type: Helm properties: ChartName: opentelemetry-collector ChartVersion: 0.127.2 ChartRepoName: opentelemetry-community ChartRepoURL: https://open-telemetry.github.io/opentelemetry-helm-charts DefaultNamespace: observability ReleaseName: opentelemetry-collector-agents - name: opentelemetry-collector-gateway description: Central gateway type: Helm properties: ChartName: opentelemetry-collector ChartVersion: 0.127.2 ChartRepoName: opentelemetry-community ChartRepoURL: https://open-telemetry.github.io/opentelemetry-helm-charts DefaultNamespace: observability ReleaseName: opentelemetry-collector-gateway ``` ## Kubernetes Manifest Amenities The `KubernetesManifest` amenity type allows you to deploy custom Kubernetes manifests to deployment cells. This is useful for deploying Secrets, ConfigMaps, or any other Kubernetes resources that are not covered by Helm chart amenities. When defining Kubernetes manifest amenities, the `properties` object supports: | Property | Type | Required | Description | | ----------- | ----- | -------- | ---------------------------------- | | `manifests` | array | Yes | List of manifest entries to deploy | Each entry in the `manifests` array must have exactly one of the following: | Property | Type | Description | | -------- | ------ | ------------------------------------------- | | `def` | object | Inline Kubernetes manifest definition | | `file` | string | Path to a YAML file containing the manifest | **Example: Inline ConfigMap** ``` customAmenities: - name: app-config description: Application configuration type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: ConfigMap metadata: name: app-config namespace: default data: DATABASE_HOST: postgres.default.svc.cluster.local LOG_LEVEL: info ``` **Example: Inline Secret** ``` customAmenities: - name: app-secrets description: Application secrets type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: Secret metadata: name: app-secrets namespace: default type: Opaque stringData: API_KEY: my-api-key ``` **Example: Docker Registry Pull Secret** When a chart depends on a private registry, it is often simpler to pre-create the pull secret as a `KubernetesManifest` amenity and then reference it from the Helm chart values. Store the base64-encoded `username:password` value in a single Omnistrate secret such as `DOCKERHUB_AUTH_B64` and inject it directly into the manifest: ``` customAmenities: - name: ingress-nginx-namespace description: Namespace for ingress components type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: Namespace metadata: name: ingress-nginx - name: ingress-nginx-pull-secret description: Pull secret for Docker Hub type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: Secret metadata: name: dockerhub-pull-secret namespace: ingress-nginx type: kubernetes.io/dockerconfigjson stringData: .dockerconfigjson: | {"auths":{"docker.io":{"auth":"$secret.DOCKERHUB_AUTH_B64"}}} ``` Tip Prefer `stringData` for Kubernetes Secrets when possible. It avoids manual base64 handling for the final manifest content and keeps the amenity definition easier to read. Composite secret transforms Amenities do not support chaining a secret lookup and another transform in the same token during manifest rendering. For example, dynamically building `base64(username:password)` from two separate secrets inside the amenity can fail on both install and uninstall. Precompute the final value into one Omnistrate secret such as `DOCKERHUB_AUTH_B64` and inject that directly. **Example: File References** ``` customAmenities: - name: external-manifests description: Manifests from files type: KubernetesManifest properties: manifests: - file: ./manifests/configmap.yaml - file: ./manifests/secret.yaml ``` File paths are resolved relative to the configuration file location. **Example: Mixed Inline and File References** ``` customAmenities: - name: mixed-manifests description: Mixed inline and file manifests type: KubernetesManifest properties: manifests: - file: ./secret.yaml - def: apiVersion: v1 kind: ConfigMap metadata: name: runtime-config namespace: default data: FEATURE_FLAG: "true" ``` **Validation** The CLI validates `KubernetesManifest` amenities before sending them to the server: - The `manifests` array cannot be empty - Each entry must have either `file` or `def`, but not both - File references must point to valid, readable YAML files - YAML content must be valid and parseable Note File references (`file`) are converted to inline definitions (`def`) by the CLI before being sent to the server API. The server only receives inline `def` entries, regardless of how they were specified in the configuration. ## Affinity rules and tolerations Omnistrate creates a system node pool that is independent of the service plan node pool. To allow Operators and Controllers to run on the system node pool, affinity rules and tolerations must be defined. Affinity rules to prevent from running on service nodes: ``` affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: omnistrate.com/control-plane operator: Exists ``` Toleration to allow to run on system nodes: ``` tolerations: - key: CriticalAddonsOnly value: "true" effect: NoSchedule ``` ## Conditional Amenities By default, every amenity in the configuration template is installed on every deployment cell. The optional `disable` field turns an amenity into a conditional amenity: an expression that is evaluated per deployment cell each time amenities are rendered. When the expression renders to `true`, the amenity is skipped on that cell; when it renders to `false`, the amenity is installed as usual. `disable` is a top-level field on both managed and custom amenities: ``` managedAmenities: - name: Kubernetes Dashboard description: Web-based Kubernetes cluster management interface type: Helm disable: $sys.deploymentCell.accountTags["skip-k8s-dashboard"] customAmenities: - name: gpu-metrics-exporter description: GPU metrics exporter type: Helm disable: $sys.deploymentCell.accountTags["skip-gpu-monitoring"] properties: # ... ``` The expression must render to `true` or `false`. A skipped amenity is also removed from the `dependsOn` lists of the remaining amenities, so dependents of a disabled amenity still install (see [Amenity Dependency Ordering](#amenity-dependency-ordering)). ### Deployment Cell Context `disable` expressions are rendered with the deployment cell's system parameters, so an amenity can be conditioned on properties of the cell it is being installed into. Useful parameters include: | Parameter | Description | | --------------------------------- | ------------------------------------------------------------------------------- | | `$sys.deploymentCell.isImported` | `true` when the cell was imported/adopted rather than provisioned by Omnistrate | | `$sys.deploymentCell.accountTags` | Custom tags of the cloud account hosting the cell (see below) | For example, to skip an amenity on adopted clusters that already run their own ingress stack: ``` managedAmenities: - name: Nginx Ingress Controller description: HTTP/HTTPS traffic routing and load balancing type: Helm disable: $sys.deploymentCell.isImported ``` ### Conditioning on Cloud Account Tags Custom tags set on a cloud account are exposed to every deployment cell running in that account as the `$sys.deploymentCell.accountTags` map. This lets you steer amenities per account — for example, tagging accounts by tier, workload type, or cost center — without maintaining a separate template per account. See [Account Tags](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#account-tags) for how to set tags with `omnistrate-ctl` and the tag semantics. The map supports the following accessor forms: | Expression | Result | | --------------------------------------------------------------- | -------------------------------------------------------------------------------------- | | `$sys.deploymentCell.accountTags` | The whole tag map as a JSON object | | `$sys.deploymentCell.accountTags.env` | Value of the `env` tag (dot notation) | | `$sys.deploymentCell.accountTags["env"]` | Value of the `env` tag (bracket notation) | | `$sys.deploymentCell.accountTags["omnistrate.com/cost-center"]` | Bracket notation for tag keys containing `.` or `/`, which dot notation cannot express | To drive a conditional amenity from a tag, set the tag value to `"true"` or `"false"` on the account and reference it directly: ``` customAmenities: - name: gpu-metrics-exporter description: GPU metrics exporter type: Helm disable: $sys.deploymentCell.accountTags["skip-gpu-monitoring"] properties: # ... ``` Referenced tags must exist A `disable` expression that references a tag key that is not present on the account does not resolve, and the amenity sync for that cell fails rather than silently installing or skipping the amenity. When you condition amenities on a tag, set that tag on **every** account (with value `"true"` or `"false"`), not only on the accounts where the amenity should be skipped. The same applies to `$sys.deploymentCell.accountTags` references in chart values: only reference tags that are guaranteed to be set. ### Using Account Tags in Chart Values The same account tags are available when amenity Helm chart values are rendered, so a single template can propagate account-level metadata into the workloads it installs: ``` customAmenities: - name: cost-exporter description: Cost and usage exporter type: Helm properties: ChartName: cost-exporter ChartVersion: 1.2.3 ChartRepoName: cost-exporter ChartRepoURL: https://charts.example.com ChartValues: commonLabels: environment: "{{ $sys.deploymentCell.accountTags.env }}" cost-center: '{{ $sys.deploymentCell.accountTags["omnistrate.com/cost-center"] }}' # Pass the whole tag map as a structured value accountTags: $sys.deploymentCell.accountTags ``` Tags can also drive derived values instead of being passed through verbatim. The `$func.if` function (see [system parameter functions](https://docs.omnistrate.com/build-guides/system-parameters/#functions)) selects between two values based on a condition, so a chart value can be computed from a tag comparison: ``` ChartValues: # "3" on accounts tagged tier=premium, "1" everywhere else replicaCount: '{{ $func.if($func.equals($sys.deploymentCell.accountTags["tier"], "premium"), "3", "1") }}' ``` ### Filtering Layered Chart Values by Account Tags Amenities also support `LayeredChartValues` — ordered value layers that are merged per deployment cell, with each layer optionally gated by a `scope`. See the [Layered Chart Values](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/index.md) guide for the full feature, including [scoped layers](https://docs.omnistrate.com/build-guides/helm-chart-layered-values/#scoped-layer) and Git-referenced layers. Scope keys are template expressions resolved against the deployment cell context, so account tags can select which layers apply on a given cell: ``` customAmenities: - name: cost-exporter description: Cost and usage exporter type: Helm properties: ChartName: cost-exporter ChartVersion: 1.2.3 ChartRepoName: cost-exporter ChartRepoURL: https://charts.example.com LayeredChartValues: # Base layer, applied on every cell - values: replicas: 1 # Applied only on cells in accounts tagged env=prod - scope: "{{ $sys.deploymentCell.accountTags.env }}": "prod" values: replicas: 3 resources: requests: cpu: 500m ``` Two behaviors to keep in mind when scoping layers on account tags: - Scope keys only support dot notation (`{{ $sys.deploymentCell.accountTags.env }}`); bracket notation is not accepted in scope keys, so tags used for layer filtering need simple keys made of letters, digits, underscores, and dots. - Unlike `disable` expressions, a scope that references a tag not set on the account does not fail the sync — the layer simply does not match and is skipped. Layers without a `scope` always apply. ## Update Configuration Template After customizing your template, apply it to your organization: ``` omnistrate-ctl deployment-cell update-config-template --cloud aws -f template-aws.yaml ``` ## View Current Configuration Template Check the current configuration template for your organization: ``` omnistrate-ctl deployment-cell describe-config-template --cloud aws ``` For JSON output: ``` omnistrate-ctl deployment-cell describe-config-template --cloud aws --output json ``` ## Managing Individual Deployment Cell Amenities ### Check Deployment Cell Status View the current amenities and status of a specific deployment cell: ``` omnistrate-ctl deployment-cell describe-config-template --id hc-xxxxxxx ``` For JSON output: ``` omnistrate-ctl deployment-cell describe-config-template --id hc-xxxxxxx --output json ``` Replace `hc-xxxxxxx` with your actual deployment cell ID. ### Update Deployment Cell Amenities You have two options for updating amenities on a deployment cell: #### Option 1: Update with Configuration File Create a custom configuration file: ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/deployment-cell-amenities-spec-schema.json managedAmenities: - name: Observability Prometheus description: Observability Prometheus type: Helm - name: Kubernetes Dashboard description: Kubernetes Dashboard type: Helm - name: External DNS description: External DNS type: Helm customAmenities: - name: Custom Chart description: Custom application chart type: Helm properties: ChartName: "my-custom-chart" ChartVersion: "1.0.0" ChartRepoURL: "https://my-repo.example.com" ``` Apply the configuration: ``` omnistrate-ctl deployment-cell update-config-template --id hc-xxxxxxx -f custom-config.yaml ``` #### Option 2: Sync with Template Automatically sync the deployment cell with your organization's template: ``` omnistrate-ctl deployment-cell update-config-template --id hc-xxxxxxx --sync-with-template ``` This command syncs the deployment cell's amenities with the configuration template defined for its environment. ### Apply Pending Changes After updating the configuration, apply the changes to trigger actual deployment: ``` omnistrate-ctl deployment-cell apply-pending-changes --id hc-xxxxxxx ``` This command: - Applies any pending amenity additions or removals - Triggers the deployment/undeployment of Helm charts - Updates the deployment cell to match the desired configuration Change Application Configuration updates create pending changes that must be explicitly applied using the `apply-pending-changes` command. This two-step process provides control and allows you to review changes before deployment. ## Amenity Dependency Ordering When one amenity depends on another resource being present first — a CRD, a namespace, a pull secret, or an operator — you can declare the relationship with `dependsOn`. Omnistrate guarantees that every listed dependency has finished installing before the dependent amenity starts. ### Why use dependsOn Without explicit ordering, all custom amenities are installed in parallel. This is fast but breaks when, for example, a Helm chart expects a CRD or a namespace that another amenity creates. `dependsOn` solves this without forcing you to bundle unrelated resources into a single chart or run multiple manual syncs. ### How it works Add a `dependsOn` list to any custom amenity entry. Each value must be the exact `name` of another amenity (custom or managed) in the same template. ``` customAmenities: - name: cert-manager description: Certificate management type: Helm properties: ChartName: cert-manager ChartVersion: v1.17.2 ChartRepoName: jetstack ChartRepoURL: https://charts.jetstack.io DefaultNamespace: cert-manager - name: cluster-issuer description: ClusterIssuer for Let's Encrypt type: KubernetesManifest dependsOn: - cert-manager properties: manifests: - def: apiVersion: cert-manager.io/v1 kind: ClusterIssuer metadata: name: letsencrypt-prod spec: acme: server: https://acme-v02.api.letsencrypt.org/directory privateKeySecretRef: name: letsencrypt-prod-key solvers: - http01: ingress: class: nginx ``` In this example, `cluster-issuer` waits for `cert-manager` to finish installing before it is applied. Without `dependsOn`, both would start simultaneously, and the ClusterIssuer would fail because the `cert-manager.io/v1` CRD does not yet exist. ### Dependency rules | Rule | Detail | | ----------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Reference by name | Each entry in `dependsOn` must match the `name` of another amenity in the same template. | | Cross-type | A Helm amenity can depend on a KubernetesManifest amenity and vice versa. | | Multiple dependencies | An amenity can list more than one dependency. It waits for all of them. | | Chains | Dependencies can be chained: A → B → C. C installs first, then B, then A. | | No cycles | Circular dependencies (A → B → A) are rejected at save time with a validation error. | | Managed amenities as targets | A custom amenity can depend on a managed amenity by name, but managed amenities cannot have `dependsOn`. This is useful to ensure a managed amenity is present and fully installed before a custom amenity starts. It also prevents accidental removal of the managed amenity while the custom amenity still needs it. | | Independent amenities stay parallel | Amenities without `dependsOn` still install in parallel. | ### Validation Omnistrate validates the dependency graph when you save the configuration template. The following errors are caught before any deployment: - **Unknown dependency**: `dependsOn` references an amenity name that does not exist in the template. - **Cyclic dependency**: The dependency graph contains a cycle (e.g., A depends on B and B depends on A). - **Managed amenity with dependsOn**: Managed amenities do not support the `dependsOn` field. ### Example: operator → CRD → application A common three-layer pattern where an operator must be running before its CRDs can be applied, and an application chart depends on those CRDs: ``` customAmenities: - name: prometheus-operator description: Prometheus Operator type: Helm properties: ChartName: kube-prometheus-stack ChartVersion: 72.6.2 ChartRepoName: prometheus-community ChartRepoURL: https://prometheus-community.github.io/helm-charts DefaultNamespace: monitoring - name: monitoring-rules description: Custom PrometheusRule and ServiceMonitor resources type: KubernetesManifest dependsOn: - prometheus-operator properties: manifests: - file: ./manifests/prometheus-rules.yaml - file: ./manifests/service-monitors.yaml - name: grafana-dashboards description: Grafana dashboard ConfigMaps type: KubernetesManifest dependsOn: - prometheus-operator properties: manifests: - file: ./manifests/grafana-dashboards.yaml ``` Here `monitoring-rules` and `grafana-dashboards` both depend on `prometheus-operator` but are independent of each other, so they install in parallel once the operator is ready. ### Example: diamond dependency ``` customAmenities: - name: shared-namespace type: KubernetesManifest properties: manifests: - def: apiVersion: v1 kind: Namespace metadata: name: data-platform - name: kafka-operator type: Helm dependsOn: - shared-namespace properties: ChartName: strimzi-kafka-operator ChartVersion: 0.45.0 ChartRepoName: strimzi ChartRepoURL: https://strimzi.io/charts/ DefaultNamespace: data-platform - name: redis-operator type: Helm dependsOn: - shared-namespace properties: ChartName: redis-operator ChartVersion: 3.5.2 ChartRepoName: ot-helm ChartRepoURL: https://ot-container-kit.github.io/helm-charts/ DefaultNamespace: data-platform - name: data-pipeline type: Helm dependsOn: - kafka-operator - redis-operator properties: ChartName: my-data-pipeline ChartVersion: 1.0.0 ChartRepoName: internal ChartRepoURL: https://charts.example.com DefaultNamespace: data-platform ``` Installation order: `shared-namespace` → `kafka-operator` + `redis-operator` (parallel) → `data-pipeline`. ### Failure and retry behavior If a dependency fails to install, all amenities that depend on it also fail. You do not need to intervene manually in partial states — re-apply pending changes and the entire amenity tree is retried from scratch: ``` omnistrate-ctl deployment-cell apply-pending-changes --id hc-xxxxxxx ``` ### Recommended patterns Use these patterns when working with amenity dependencies: - If strict ordering is required and the resources are tightly coupled, prefer a single Helm chart that owns both the prerequisite objects and the application resources. - If the resources have separate lifecycles (different upgrade cadences, different owners), use `dependsOn` to declare the ordering between separate amenities. - If several charts share the same registry credentials, keep the namespace and image pull secret as dedicated manifest amenities with `dependsOn` pointing from the charts to the secrets. - If the chart can create its own namespace, use `DefaultNamespace` and chart-native secret creation rather than a separate namespace amenity. ## Debugging Amenities Sync When an amenities rollout is stuck or a new template cannot be applied because a sync is still in progress, inspect the active deployment-cell workflows first: ``` omnistrate-ctl deployment-cell workflow list hc-xxxxxxx omnistrate-ctl deployment-cell workflow events hc-xxxxxxx ``` These commands show whether the cell is still in `DEPLOYMENT_CELL_SYNC_AMENITIES`, which step is currently active, and the last debug event emitted by the workflow. If a previous amenities sync is blocking a new update: - If the sync appears as its own deployment-cell workflow in the UI, cancel that workflow there and rerun the update. - If the cell sync was triggered indirectly by another workflow, inspect the deployment-cell workflow list first. The parent workflow can finish or fail while the cell sync is still active. - If the workflow remains stuck and no cancel control is available, wait for the workflow timeout or contact support rather than layering more template changes on top of the same stuck sync. If the workflow says a Helm amenity failed, connect to the deployment cell and inspect the releases directly: ``` omnistrate-ctl deployment-cell update-kubeconfig hc-xxxxxxx --role cluster-admin --kubeconfig /tmp/kubeconfig export KUBECONFIG=/tmp/kubeconfig helm ls -A kubectl get pods -A ``` Use this flow when: - A sync is still running even though the releases appear installed - A reverted template still fails because a manifest or release is not reconciling cleanly - A new `update-config-template` attempt is blocked by an older sync that is still in progress ## Cancel and Restart Amenity Operations If an amenity deployment or update operation is taking too long or has encountered issues, you can cancel the operation and restart it. ### Restart a failed amenity operation If an amenity operation has failed, you can restart it by re-applying the pending changes: ``` omnistrate-ctl deployment-cell apply-pending-changes --id hc-xxxxxxx ``` This retries the deployment of any amenities that failed during the previous attempt. ## Per-Environment Deployment Cells Omnistrate supports configuring deployment cells on a per-environment basis. This allows you to have different deployment cell configurations and amenities for each environment (e.g., Dev, Staging, Production). ### How it works When you create deployment cells for different environments, each environment maintains its own set of deployment cells with independent configurations: - **Dev environment**: Minimal amenities, smaller node pools - **Staging environment**: Production-like amenities for testing - **Production environment**: Full amenity suite with high-availability configurations ### Configuring amenities per environment Each environment's deployment cells can have their own configuration templates. To set up environment-specific amenities: 1. Generate a configuration template for the target environment's deployment cell. 1. Customize the amenities for that environment. 1. Apply the configuration to the deployment cell. ``` # List deployment cells omnistrate-ctl deployment-cell list # Generate a configuration template for a cloud provider omnistrate-ctl deployment-cell generate-config-template --cloud aws # Update amenities for a specific deployment cell omnistrate-ctl deployment-cell update-config-template --id hc-xxxxxxx -f config.yaml # Apply changes omnistrate-ctl deployment-cell apply-pending-changes --id hc-xxxxxxx ``` Tip Use configuration templates to maintain consistent amenity configurations across deployment cells within the same environment. Generate a template first to see the available options, then customize and apply it. # Deployment Cell Node Pool Management The `deployment-cell` CTL command now includes comprehensive node pool management capabilities for AWS, GCP, Azure, and OCI deployment cells. For more details, see the CTL documentation: [Deployment Cells](https://ctl.omnistrate.cloud/omnistrate-ctl_deployment-cell/). ## Why AWS cells can show many nodegroups On AWS, Omnistrate may create distinct managed nodegroups for each effective combination of instance type, availability zone, and placement mode. Over time, especially after compute iteration, a cell can accumulate zero-sized nodegroups that are no longer hosting workload but still count toward EKS nodegroup quotas. Use this page to identify, scale down, or delete stale nodegroups. For quota planning and quota increase recommendations, see [Service Quotas](https://docs.omnistrate.com/runtime-guides/cloud-provider-quotas/index.md). ## When to Use Node Pool Management ### CUSTOM_TENANCY Plans (Manual Management Required) Node pool management commands are **primarily designed for CUSTOM_TENANCY plans** where workloads are deployed through Helm, Operators, or Kustomize. In these scenarios, you may need to manually manage node pools to: - **Scale down node pools** to reduce infrastructure costs during off-hours or low-usage periods - **Clean up stale or inactive node pools** that are no longer needed but consuming resources - **Optimize costs** by removing node pools associated with terminated or inactive customer instances These commands provide a single-command way to manage node pool capacity and cleanup, helping you control infrastructure costs effectively. ### OMNISTRATE_DEDICATED_TENANCY and OMNISTRATE_MULTI_TENANCY Plans (Automatic Management) For plans using **OMNISTRATE_DEDICATED_TENANCY** or **OMNISTRATE_MULTI_TENANCY** with resources imported through Docker Compose / Compose Spec, the platform **automatically manages node pool lifecycle**: - Automatic scale-up and scale-down based on workload demand - Automatic cleanup of node pools when instances are terminated - No manual intervention required for cost optimization You typically do **not need** to use these commands for automatically managed plans unless you have specific infrastructure management requirements. ## Commands Overview ### list-nodepools List all node pools in a deployment cell with their configuration details. **Usage:** ``` omnistrate-ctl deployment-cell list-nodepools --id ``` **Options:** - `--id, -i`: Deployment cell ID (required) - `--output, -o`: Output format - `table` (default), `text`, or `json` **Output Fields:** - **Name**: Node pool identifier - **Type**: Entity type (NODEPOOL, NODE_GROUP, or AZURE_NODEPOOL) - **MachineType**: Instance/VM type (e.g., n2-highmem-2, t3.medium, Standard_E2as_v6) - **ImageType**: OS image type (e.g., COS_CONTAINERD, AL2_x86_64) - **MinNodes**: Minimum autoscaling node count - **MaxNodes**: Maximum autoscaling node count - **Location**: Zone, availability zone, or subnet location - **AutoRepair**: Automatic node repair enabled (GCP) - **AutoUpgrade**: Automatic node upgrade enabled (GCP) - **AutoScaling**: Autoscaling enabled (Azure) - **CapacityType**: Node capacity type - ON_DEMAND or SPOT (AWS) - **PrivateSubnet**: Whether nodes use private subnets **Examples:** ``` # List all nodepools in table format omnistrate-ctl deployment-cell list-nodepools --id hc-9c5ok6tmv # List nodepools as JSON omnistrate-ctl deployment-cell list-nodepools --id hc-9c5ok6tmv -o json ``` ______________________________________________________________________ ### describe-nodepool Get detailed information about a specific node pool, including current node count. **Usage:** ``` omnistrate-ctl deployment-cell describe-nodepool --id --nodepool ``` **Options:** - `--id, -i`: Deployment cell ID (required) - `--nodepool, -n`: Node pool name (required) - `--output, -o`: Output format - `table` (default), `text`, or `json` **Output Fields:** All fields from `list-nodepools` plus: - **CurrentNodes**: Current number of running nodes in the pool **Examples:** ``` # Describe a GCP nodepool omnistrate-ctl deployment-cell describe-nodepool \ --id hc-9c5ok6tmv \ --nodepool pt-uzdahfq76b-n2-highmem-2-a # Describe an AWS nodegroup with JSON output omnistrate-ctl deployment-cell describe-nodepool \ --id hc-9sp0n4418 \ --nodepool hc-9sp0n4418-pt-uzdahfq76b-r7i-large-us-east-1c \ -o json # Describe an Azure nodepool omnistrate-ctl deployment-cell describe-nodepool \ --id hc-nnqkjzz9j \ --nodepool izahdemsjqy9 ``` ______________________________________________________________________ ### scale-down-nodepool Scale down a node pool to zero nodes for cost savings. **Usage:** ``` omnistrate-ctl deployment-cell scale-down-nodepool --id --nodepool ``` **Options:** - `--id, -i`: Deployment cell ID (required) - `--nodepool, -n`: Node pool name (required) **Behavior:** - Sets the node pool's maximum size to 0 - Evicts all running nodes in the pool - Node pool configuration remains intact - Can be reversed with `scale-up-nodepool` - Useful for reducing costs during off-hours or low-usage periods **Examples:** ``` # Scale down a GCP nodepool omnistrate-ctl deployment-cell scale-down-nodepool \ --id hc-9c5ok6tmv \ --nodepool pt-uzdahfq76b-n2-highmem-2-a # Scale down an AWS nodegroup omnistrate-ctl deployment-cell scale-down-nodepool \ --id hc-9sp0n4418 \ --nodepool hc-9sp0n4418-r-4qouebzi1o-t3-medium-us-east-1c ``` ______________________________________________________________________ ### scale-up-nodepool Restore a node pool to its default maximum capacity of 450 nodes. **Usage:** ``` omnistrate-ctl deployment-cell scale-up-nodepool --id --nodepool ``` **Options:** - `--id, -i`: Deployment cell ID (required) - `--nodepool, -n`: Node pool name (required) **Behavior:** - Sets the node pool's maximum size to 450 (default for all clouds) - Restores autoscaling capacity after scale-down - Nodes are provisioned on-demand by the autoscaler as workloads require - Does not immediately create nodes - autoscaler manages actual node count **Examples:** ``` # Scale up a previously scaled-down nodepool omnistrate-ctl deployment-cell scale-up-nodepool \ --id hc-9c5ok6tmv \ --nodepool pt-uzdahfq76b-n2-highmem-2-a # Restore an AWS nodegroup to default capacity omnistrate-ctl deployment-cell scale-up-nodepool \ --id hc-9sp0n4418 \ --nodepool hc-9sp0n4418-pt-uzdahfq76b-r7i-large-us-east-1c ``` ______________________________________________________________________ ### delete-nodepool Permanently delete a node pool from a deployment cell. **Usage:** ``` omnistrate-ctl deployment-cell delete-nodepool --id --nodepool ``` **Options:** - `--id, -i`: Deployment cell ID (required) - `--nodepool, -n`: Node pool name (required) **Behavior:** - Permanently deletes the node pool configuration - Evicts all nodes and removes the pool from the cluster - Operation can take up to 10 minutes - Shows a spinner during deletion - Cannot be reversed - use `scale-down-nodepool` if you want to preserve the configuration **Examples:** ``` # Delete a GCP nodepool omnistrate-ctl deployment-cell delete-nodepool \ --id hc-9c5ok6tmv \ --nodepool pt-uzdahfq76b-n2-highmem-2-a # Delete an AWS nodegroup omnistrate-ctl deployment-cell delete-nodepool \ --id hc-9sp0n4418 \ --nodepool hc-9sp0n4418-r-4qouebzi1o-t3-medium-us-east-1c ``` ______________________________________________________________________ ## Node Pool Cleanup Certain configuration changes — such as switching the OS family, changing instance types, or modifying placement constraints — cause Omnistrate to create new node pools. The old node pools are not deleted immediately, ensuring a safe rollover period while workloads migrate to the new nodes. After verifying that workloads are running on the new node pools, use the commands documented above ([list-nodepools](#list-nodepools), [delete-nodepool](#delete-nodepool)) to identify and remove stale node pools. Check node pool quota before making changes AWS EKS enforces a limit on the number of managed node groups per cluster. Before triggering changes that create new node pools, verify that your cluster has sufficient quota. For quota details and increase recommendations, see [Service Quotas](https://docs.omnistrate.com/runtime-guides/cloud-provider-quotas/index.md). ## Common Workflows ### Cost Optimization ``` # Scale down nodepools during off-hours omnistrate-ctl deployment-cell scale-down-nodepool --id hc-xyz --nodepool my-nodepool # Scale up when traffic returns omnistrate-ctl deployment-cell scale-up-nodepool --id hc-xyz --nodepool my-nodepool ``` ### Node Pool Lifecycle ``` # 1. View existing nodepools omnistrate-ctl deployment-cell list-nodepools --id hc-xyz # 2. Get details on a specific node pool to get the current node count omnistrate-ctl deployment-cell describe-nodepool --id hc-xyz --nodepool my-nodepool # 3. Scale down for cost savings omnistrate-ctl deployment-cell scale-down-nodepool --id hc-xyz --nodepool my-nodepool # 4. Later, restore capacity omnistrate-ctl deployment-cell scale-up-nodepool --id hc-xyz --nodepool my-nodepool # 5. Or permanently remove if no longer needed omnistrate-ctl deployment-cell delete-nodepool --id hc-xyz --nodepool my-nodepool ``` # Deployment Instances ## What are Deployment Instances? Deployment instances are running instances of your SaaS product that have been provisioned for your customers. Key characteristics of deployment instances: - **Tenant-specific deployments**: Each instance serves specific tenant and their data - **Resource allocation**: Instances have dedicated compute, storage, and networking resources - **Lifecycle management**: Instances can be started, stopped, scaled, updated, and deleted - **Environment isolation**: Instances are deployed within specific environments (Dev, Staging, Production) - **Plan-based configuration**: Each instance is created based on a specific plan that defines its resources and capabilities Deployment instances are the core operational units of your SaaS product - they represent the actual running services that your customers interact with. ## Managing Deployment Instances You can manage deployment instances through multiple interfaces: - **Customer Portal**: Self-serve [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md) for customer to manage their deployment instances - **Operations Console UI**: The Deployment Instances page provides a visual interface for monitoring and managing your customer instances - **CLI**: Command-line tools allow for programmatic management and automation - **API**: RESTful APIs enable integration with your own tools and workflows ## Manual Operations You can perform various manual operations on deployment instances through both the web interface and CLI/API: ### Stopping Instances You can manually stop running deployment instances when needed for maintenance, troubleshooting, or resource management. This operation gracefully shuts down the instance while preserving its configuration. ### Deleting Instances When instances are no longer needed, you can permanently delete them to free up resources and clean up your deployment inventory. This operation removes the instance and its associated resources. ### Restarting Instances Restart instances to apply updates or resolve issues. This operation provides a clean restart of the instance while maintaining its configuration and data. ### Scaling Instances Adjust resource allocation based on demand. This allows you to increase or decrease compute, memory, and storage resources as needed. ### Updating Configurations Modify instance settings without full redeployment. This enables you to change configuration parameters while keeping the instance running. All these operations can be performed through the web interface for manual management or via CLI/API for automated workflows. These operations can be also enabled in the Customer Portal to allow your customers to perform these actions. ### Changing Network Type You can change the network type of a running deployment instance between **public** and **private** endpoints. This allows you to adjust network accessibility after deployment without recreating the instance. #### When to change network type Common scenarios for changing network type include: - **Moving to private**: Transitioning from a public endpoint to a private one for improved security, compliance requirements, or when customers establish [VPC peering](https://docs.omnistrate.com/tenant-management/private-networking/index.md) or [PrivateLink](https://docs.omnistrate.com/tenant-management/private-networking/#aws-privatelink) connectivity. - **Moving to public**: Temporarily exposing an instance publicly for debugging, onboarding, or when private connectivity is not yet configured. #### How to change network type To change the network type of a deployment instance: 1. Navigate to the deployment instance in the Operations Console. 1. Select **Update** on the instance. 1. Change the **Network Type** parameter from public to private (or vice versa). 1. Confirm the operation. The network type change runs as a [workflow](https://docs.omnistrate.com/operate-guides/workflows/index.md), so you can track its progress and troubleshoot any issues. Warning Changing the network type updates the instance endpoint. Any clients connecting to the instance must be updated to use the new endpoint after the change completes. Coordinate with your customers before making this change. Note Private network types require that the customer has established private connectivity (such as VPC peering or PrivateLink) to reach the instance. Verify connectivity is in place before switching to a private endpoint. ## Instance Tags Instance tags allow you to organize and categorize your deployment instances with custom key-value pairs. Tags are useful for filtering, grouping, and managing instances across your fleet. ### Adding tags to an instance You can add tags to an instance during creation or update them on an existing instance: 1. Navigate to **Manage Fleet > Instances**. 1. Select the instance you want to tag. 1. Click **Edit Tags** to add, modify, or remove tags. Tags are key-value pairs where: - **Key**: A string identifier for the tag (e.g., `team`, `project`, `cost-center`) - **Value**: The value associated with the tag key (e.g., `backend`, `payments`, `engineering`) ### Filtering by tags Use tags to filter instances in the Operations Console: 1. Navigate to **Manage Fleet > Instances**. 1. Use the tag filter to select specific tag keys and values. 1. The instance list updates to show only matching instances. Tip Establish a consistent tagging strategy across your organization to make fleet management easier. Common tags include `team`, `project`, `region`, and `cost-center`. # Deployment Snapshots ## What are Deployment Snapshots? Deployment snapshots are **manual, on-demand snapshots** of a deployment instance's persistent storage that you can use for disaster recovery. Snapshots capture the state of all backed-up volumes at the time of creation. Key characteristics of deployment snapshots: - **Manual triggers**: Create snapshots on-demand as part of your disaster recovery preparedness and runbooks. - **Cross-region disaster recovery**: Copy snapshots to a different region than the source deployment. This ensures that your data remains available even in the event of a total regional cloud outage. - **Independent lifecycle**: Unlike automated backups that are cleaned up automatically, manual snapshots are **persistent**. They remain in your account until you explicitly delete them. - **Multi-cloud support**: Snapshots are supported on AWS, GCP, and Azure across all storage types that support backups (EBS, Persistent Disk, Azure Disk). Deployment snapshots are most useful when you want a durable recovery artifact for disaster recovery that you control, separate from automated schedules and retention policies. For automated backups and point-in-time recovery, see [Backup and Point-in-time Restore (PITR)](https://docs.omnistrate.com/runtime-guides/pitr/index.md). ## Cloud Provider Support | Capability | AWS | GCP | Azure | | -------------------- | --- | --- | ----- | | Create snapshot | | | | | Same-region restore | | | | | Cross-region copy | | | | | Cross-region restore | | | | Note Snapshots are available for resources that use backed-up persistent storage (EBS, GCP Persistent Disk, Azure Disk). Volumes with `disableBackup: true` in their storage configuration are excluded from snapshots. ## Managing Deployment Snapshots You can manage deployment snapshots through multiple interfaces: - **Operations Console UI**: Access snapshot management from two levels: - **Deployment Snapshots page**: View and manage snapshots across a service environment. - **Instance Snapshots tab**: From a deployment instance details view, manage snapshots for that specific instance. - **Customer Portal**: If enabled, your customers can create and manage snapshots from the [Customer Portal](https://docs.omnistrate.com/tenant-management/customer-portal/index.md). - **API/CLI**: Use the Omnistrate API or [CTL](https://docs.omnistrate.com/getting-started/installing-ctl/index.md) for programmatic snapshot management and automation. For operational visibility (including progress and troubleshooting), see [Workflows](https://docs.omnistrate.com/operate-guides/workflows/index.md). ## Supported Operations Deployment snapshots support the following operations: | Operation | Description | | ----------- | ------------------------------------------------------------ | | **Create** | Capture the current state of a deployment instance's volumes | | **Copy** | Replicate a snapshot to a different region | | **Restore** | Recover a deployment instance from a snapshot | | **Delete** | Permanently remove a snapshot | All snapshot operations run as [workflows](https://docs.omnistrate.com/operate-guides/workflows/index.md), so you can track progress and troubleshoot issues through the workflow details view. In cross-region scenarios, both snapshot creation and snapshot copy can incur inter-region transfer and storage costs. ### Create a snapshot Create a snapshot when you want a durable recovery artifact for a deployment instance. Snapshots remain available until you explicitly delete them. To create a snapshot: 1. Navigate to the deployment instance in the Operations Console. 1. Open the **Snapshots** tab. 1. Click **Create Snapshot**. 1. Optionally, select a **destination region** for cross-region disaster recovery. If you don't select a destination region, the snapshot is created in the same region as the instance. 1. Confirm the operation. Warning Snapshot creation locks the deployment instance status and blocks other concurrent operations while in progress. Plan a maintenance window and communicate the expected impact to your customers. ### Copy a snapshot Copy a snapshot to make it available in another region. After copying, you can restore from the copied snapshot in the destination region. This is the foundation for cross-region disaster recovery. To copy a snapshot: 1. Navigate to the **Deployment Snapshots** page or the instance's **Snapshots** tab. 1. Select the snapshot you want to copy. 1. Click **Copy** and select the **destination region**. 1. Confirm the operation. Warning The destination region must be supported by the deployment instance's Plan. If the region is not supported, the copy fails. Cross-region copies incur network transfer costs. ### Restore from a snapshot Restore recovers a deployment instance to the state captured by a specific snapshot. Restore is only possible in the **same region** as the snapshot. To restore from a snapshot: 1. Navigate to the snapshot you want to restore from. 1. Click **Restore**. 1. Confirm the operation. If you need to restore in a different region, copy the snapshot to that region first, then restore from the copy. Tip Before restoring, consider creating a new snapshot of the current state as a safety measure in case you need to revert. ### Delete a snapshot Delete a snapshot when you no longer need it. Because snapshots are persistent by design, periodically review and delete snapshots you no longer need to reduce storage costs. Deletion is **permanent**. After you delete a snapshot, you cannot restore from it. Warning Deleting a snapshot removes only that specific snapshot artifact. If you copied the snapshot to other regions, you must delete those copies separately. ## Cross-Region Disaster Recovery Snapshots enable cross-region disaster recovery on all supported cloud providers (AWS, GCP, and Azure). Use this workflow to prepare for regional outages: 1. **Create** a snapshot of the deployment instance in its source region. 1. **Copy** the snapshot to a target region in a different geographic area. 1. In a disaster scenario, **restore** from the copied snapshot in the target region. Best practices for cross-region DR - Copy snapshots to at least one geographically distant region. - Establish a regular cadence for creating snapshots (for example, before major changes or upgrades). - Test your restore process periodically to ensure your recovery time objectives are achievable. - Monitor snapshot operations through [Alarms](https://docs.omnistrate.com/operate-guides/alarms/index.md) — events like `FailedSnapshotCopy` and `SuccessfulSnapshotCreate` help you track the health of your DR process. ## Snapshot-on-Delete You can configure your SaaS Product to automatically create a final snapshot before a deployment instance is deleted. This provides a safety net against accidental data loss. To enable this behavior, set `snapshotBeforeDeletion: true` in the backup configuration of your service specification: === "Docker Compose" ``` x-omnistrate-capabilities: backupConfiguration: backupRetentionInDays: 7 backupPeriodInHours: 24 snapshotBeforeDeletion: true ``` For more details, see the [Backup Configuration](https://docs.omnistrate.com/build-guides/compose-spec/#x-omnistrate-capabilitiesbackupconfiguration) section in the Compose specification. ## Monitoring Snapshot Operations Omnistrate provides alarm events for snapshot operations. You can subscribe to these events to monitor the health of your snapshot workflows: | Event | Description | | -------------------------- | ---------------------------------------------------- | | `SuccessfulSnapshotCreate` | A snapshot was created successfully | | `FailedSnapshotCreate` | A snapshot creation failed | | `SuccessfulSnapshotCopy` | A snapshot was copied to another region successfully | | `FailedSnapshotCopy` | A snapshot copy failed | | `SuccessfulSnapshotDelete` | A snapshot was deleted successfully | | `FailedSnapshotDelete` | A snapshot deletion failed | For details on configuring alarm notifications, see [Alarms](https://docs.omnistrate.com/operate-guides/alarms/index.md). # Fleet Dashboard The Fleet Dashboard provides comprehensive visibility into your SaaS fleet operations, offering real-time insights into deployment health, workflow status, patching progress, and operational metrics across all your deployment models and cloud environments. ## Overview Managing a distributed SaaS fleet across multiple cloud providers and deployment models presents unique operational challenges. The Fleet Dashboard serves as your central command center, providing unified visibility into: - **Fleet Health**: Real-time status of all deployment instances - **Workflow Management**: Active, completed, and failed operations - **Patching Operations**: Version updates and maintenance activities - **Alert Management**: System events and notifications - **Resource Inventory**: Plans, subscriptions, and deployments ## Why Fleet Visibility Matters ### Operational Complexity at Scale As your SaaS grows, operational complexity increases exponentially. Without proper visibility, you face: - **Blind Spots**: Inability to detect issues before they impact customers - **Reactive Operations**: Discovering problems only after customer complaints - **Resource Waste**: Inefficient resource allocation and utilization - **Compliance Risks**: Difficulty tracking changes and maintaining audit trails - **Scaling Challenges**: Manual processes that don't scale with growth ### Multi-Cloud and Multi-Deployment Model Challenges Operating across different cloud providers and deployment models introduces additional complexity: **Cloud Provider Variations** - Different APIs, tools, and operational paradigms - Varying service availability and feature sets across regions - Inconsistent monitoring and alerting capabilities - Diverse cost structures and billing models **Deployment Model Diversity** - **Hosted SaaS**: Your infrastructure, your responsibility - **BYOC (Bring Your Own Cloud)**: Customer infrastructure, shared responsibility - **Air-Gapped**: Off-line instalations, limited access Each model requires different operational approaches while maintaining consistent service quality. ## Dashboard Components ### Fleet Overview Metrics The dashboard provides key performance indicators at a glance: - **Plans**: Available service tiers and configurations - **Workflows**: Active operational processes - **Subscriptions**: Customer subscription status - **Deployment Instances**: Running service instances - **Deployment Cells**: Logical groupings of resources ### Health Summary Monitor the overall health of your fleet: - **Total Instances**: Complete inventory of deployment instances - **Healthy Instances**: Instances operating within normal parameters - **Unhealthy Instances**: Instances requiring attention or intervention - **Unknown Instances**: Instances that are not yet reporting health status ### Workflow Summary Track operational activities across your fleet: - **Active Workflows**: Currently running operations - **Completed Workflows**: Successfully finished operations - **Failed Workflows**: Operations requiring investigation or retry ### Patching Overview Monitor version updates and maintenance activities: - **Version Distribution**: Current software versions across the fleet - **Patching Progress**: Real-time status of ongoing updates - **Completion Rates**: Success metrics for maintenance operations ## Alert Management The dashboard includes comprehensive alerting capabilities. For detailed alert configuration, see [Alert Management](https://docs.omnistrate.com/operate-guides/alarms/index.md). # Instance Breakpoints Instance breakpoints let you pause a **create** workflow **before** a target resource is reconciled, inspect state, and then resume execution. ## What this supports - Breakpoints are currently supported for **instance create** workflows. - You can define breakpoints by **resource key** or **resource ID**. - You can optionally pause on supported resource events by using `:|`. - When a breakpoint is hit, the workflow is paused until resumed. - Breakpoint status is visible in: - `omctl instance breakpoint list` - `omctl instance debug` (interactive TUI and JSON) - Breakpoints are deleted after the workflow completes. ## Resource events Without an event, a breakpoint pauses before the target resource is reconciled. Event-specific breakpoints pause at supported lifecycle points inside a resource reconciliation. | Resource type | Supported events | | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | Helm | `StartHelmInstall`, `CompleteHelmInstall`, `FailHelmInstall` | | Terraform | `StartTerraformPlan`, `CompleteTerraformPlan`, `FailTerraformPlan`, `StartTerraformApply`, `CompleteTerraformApply`, `FailTerraformApply` | | Other resource types | No resource events | Event-specific breakpoints are checked in the resource reconciliation timeline: If a configured breakpoint matches one of these event names, Omnistrate pauses at that point before moving to the next step. ## Example setup Use the following full Service Plan spec and save it as `./instance-breakpoints-spec.yaml`: ``` # yaml-language-server: $schema=https://api.omnistrate.cloud/2022-09-01-00/schema/service-spec-schema.json services: - name: serviceA description: Service A passive: true dependsOn: - serviceB - name: serviceB description: Service B passive: true ``` In this spec, `serviceA` depends on `serviceB`, so `serviceB` is created first. ## 1. Build and release product ``` omctl build -f ./instance-breakpoints-spec.yaml \ --product-name 'TestBreakpoints' \ --spec-type ServicePlanSpec \ --release-as-preferred ``` ## 2. Create instance with breakpoint Set a breakpoint on `serviceA` during create. Because `serviceA` depends on `serviceB`, the workflow reconciles `serviceB` first and then pauses before reconciling `serviceA`: ``` omctl instance create \ --cloud-provider aws \ --environment dev \ --service TestBreakpoints \ --plan TestBreakpoints \ --region us-east-2 \ --resource serviceA \ --breakpoints serviceA ``` To pause on Helm install events, include the event names after the resource key or ID: ``` omctl instance create \ --cloud-provider aws \ --environment dev \ --service TestBreakpoints \ --plan TestBreakpoints \ --region us-east-2 \ --resource serviceA \ --breakpoints chart:StartHelmInstall|FailHelmInstall ``` For Terraform resources: ``` omctl instance create \ --cloud-provider aws \ --environment dev \ --service TestBreakpoints \ --plan TestBreakpoints \ --region us-east-2 \ --resource serviceA \ --breakpoints tf:StartTerraformPlan|CompleteTerraformPlan|FailTerraformPlan|StartTerraformApply|CompleteTerraformApply|FailTerraformApply ``` ## 3. List active breakpoints Text output includes both resource **id** and **key**: ``` omctl instance breakpoint list instance-ckxvekao9 ``` ``` ╭────────────────────────────────────────────╮ │ id key event status │ │────────────────────────────────────────────│ │ r-Tz8WmKj4Xn serviceA StartHelmInstall Hit │ ╰────────────────────────────────────────────╯ ``` JSON output: ``` omctl instance breakpoint list instance-ckxvekao9 -o json ``` ``` [ { "key": "serviceA", "id": "r-Tz8WmKj4Xn", "event": "StartHelmInstall", "status": "Hit" } ] ``` ## 4. Inspect paused workflow with debug ### Interactive TUI ``` omctl instance debug instance-ckxvekao9 ``` When a breakpoint is hit: - Header shows workflow ID with `PAUSED`. - Target resource card shows `BREAKPOINT` and `breakpoint hit`. ### JSON view ``` omctl instance debug instance-ckxvekao9 -o json ``` Breakpoint status is available under: - `planDag.breakpointById` - `planDag.breakpointByKey` - `planDag.breakpointByName` - `planDag.breakpointEventsById` - `planDag.breakpointEventsByKey` - `planDag.breakpointEventsByName` Example: ``` { "instanceId": "instance-ckxvekao9", "planDag": { "workflowId": "submit-create-instance-ckxvekao9-1773256161588972", "breakpointById": { "r-Tz8WmKj4Xn": "hit" }, "breakpointByKey": { "serviceA": "hit" }, "breakpointByName": { "serviceA": "hit" }, "breakpointEventsByKey": { "serviceA": { "StartHelmInstall": "hit" } } } } ``` ## 5. Resume paused workflow Resume with confirmation: ``` omctl instance breakpoint resume instance-ckxvekao9 ``` Confirmation includes both resource key and ID: ``` ┃ Resume workflow submit-create-instance-ckxvekao9-1773256161588972 for instance instance-ckxvekao9? ┃ Breakpoint to resume: serviceA [r-Tz8WmKj4Xn] ``` Event-specific breakpoints include the event: ``` ┃ Breakpoint to resume: serviceA [r-Tz8WmKj4Xn] @ StartHelmInstall ``` Skip confirmation: ``` omctl instance breakpoint resume instance-ckxvekao9 --force ``` ## 6. Verify workflow completion After resume, debug shows completed progress and no breakpoint maps: ``` omctl instance debug instance-ckxvekao9 -o json ``` Expected behavior: - `progressById`, `progressByKey`, and `progressByName` show `completed`. - `breakpointById` / `breakpointByKey` / `breakpointByName` are no longer present. You can also re-run breakpoint list to verify no active breakpoints remain: ``` omctl instance breakpoint list instance-ckxvekao9 ``` ## Related commands - `omctl instance create --breakpoints [:[|...]][,...]` - `omctl instance breakpoint list ` - `omctl instance breakpoint resume [--force]` - `omctl workflow resume ` # Managed Workload Identities ## Overview Helm amenities or service workloads might require permissions necessary to call cloud-native services outside the Kubernetes cluster. For example, a custom component might need to publish messages to an external queue, write objects to a bucket, or access a managed database. The managed workload identities feature allows you to define cloud permissions that your workloads need and use the resulting cloud identity from the Kubernetes service account. Omnistrate will provision all required cloud provider-specific entities during account onboarding and will create federated access so the workloads can use such external identity. The identity and federation workflow is cloud-specific. You define the permissions in the deployment cell configuration, and Omnistrate handles the provider-specific role, service account, trust, and federation details required for that cloud. ## Cloud Providers The following sections describe the managed workload identity flow for each supported cloud provider. The AWS flow is documented first; the other provider sections will be expanded as their examples are added. ### AWS #### Create an AWS managed workload identity Define the identity in the `managedIdentities` section of the AWS amenities template. The example below creates a `queue-writer` identity with permission to send messages only to SQS queues whose names begin with `orders-`. Service account binding specifies which Kubernetes service account will be allowed to use the identity: ``` managedIdentities: - identifier: queue-writer description: Allows the workload to publish messages to application queues. bindings: - serviceAccount: namespace: queue-system name: queue-writer permissions: policies: aws: | { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "sqs:SendMessage" ], "Resource": "arn:aws:sqs:*:*:orders-*" } ] } ``` When the account is onboarded, Omnistrate creates the AWS IAM role and its policy for this managed workload identity in the customer account. You do not need to create the role manually. Omnistrate also configures the AWS federation needed for Kubernetes workloads to use the role during deployment cell provisioning. #### Resolve the ARN of the managed identity role in the system parameters At deployment time, Omnistrate exposes the role ARN through the [`$sys.deploymentCell.aws.managedWorkloadIdentities["managedIdentityIdentifier"].roleARN`](https://docs.omnistrate.com/build-guides/system-parameters/#aws-specific-deployment-cell-parameters) system parameter. Replace `managedIdentityIdentifier` with the identity's `identifier` value. For the example above, use `$sys.deploymentCell.aws.managedWorkloadIdentities["queue-writer"].roleARN`. The value is resolved for the target deployment cell, so the manifest does not need to contain a hard-coded AWS account ID or role ARN. #### Use the role from a Kubernetes service account Annotate the service account used by the workload with the resolved role ARN. For Amazon EKS, the `eks.amazonaws.com/role-arn` annotation associates the service account with the IAM role: ``` apiVersion: v1 kind: ServiceAccount metadata: name: queue-writer namespace: application annotations: eks.amazonaws.com/role-arn: '{{ $sys.deploymentCell.aws.managedWorkloadIdentities["queue-writer"].roleARN }}' ``` Such a service account can be defined in the amenity template or a service plan spec (for use by service instance). ### GCP #### Create a GCP managed workload identity Define the identity in the `managedIdentities` section of the deployment cell configuration. The example below creates a `queue-writer` identity with the `roles/pubsub.publisher` role. Service account bindings specify which Kubernetes service accounts are allowed to use the identity: ``` managedIdentities: - identifier: queue-writer description: Allows the workload to publish messages to Pub/Sub topics. bindings: - serviceAccount: namespace: queue-writer name: queue-writer permissions: roles: gcp: - name: roles/pubsub.publisher ``` When the account is onboarded, the account setup script creates the GCP service account for this managed workload identity in the customer project. Omnistrate configures the GKE Workload Identity federation needed for Kubernetes workloads to authenticate as the service account during deployment cell provisioning. #### Resolve the service account email in the system parameters At deployment time, Omnistrate exposes the service account email through the [`$sys.deploymentCell.gcp.managedWorkloadIdentities["managedIdentityIdentifier"].serviceAccountEmail`](https://docs.omnistrate.com/build-guides/system-parameters/#gcp-specific-deployment-cell-parameters) system parameter. Replace `managedIdentityIdentifier` with the identity's `identifier` value. For the example above, use `$sys.deploymentCell.gcp.managedWorkloadIdentities["queue-writer"].serviceAccountEmail`. The value is resolved for the target deployment cell, so the manifest does not need to contain a hard-coded project ID or service account name. #### Use the identity from a Kubernetes service account Annotate the service account used by the workload with the resolved GCP service account email. For GKE Workload Identity, the `iam.gke.io/gcp-service-account` annotation associates the Kubernetes service account with the GCP service account: ``` apiVersion: v1 kind: ServiceAccount metadata: name: queue-writer namespace: queue-writer annotations: iam.gke.io/gcp-service-account: '{{ $sys.deploymentCell.gcp.managedWorkloadIdentities["queue-writer"].serviceAccountEmail }}' ``` You can define such a service account in an amenity template or a plan spec (for use by a product instance). ### Azure #### Create an Azure managed workload identity Define the identity in the `managedIdentities` section of the deployment cell configuration. The example below creates a `blob-writer` identity with the `Storage Blob Data Contributor` role. Service account bindings specify which Kubernetes service accounts are allowed to use the identity: ``` managedIdentities: - identifier: blob-writer description: Read, write, and delete Azure Storage containers and blobs. bindings: - serviceAccount: namespace: blob-writer-ns name: blob-writer permissions: roles: azure: - name: "Storage Blob Data Contributor" ``` When the account is onboarded, the account setup script creates the Azure application registration and associated service principal for this managed workload identity in the customer's Entra ID tenant. Omnistrate configures the federated identity credential needed for Kubernetes workloads to authenticate as the service principal during deployment cell provisioning. #### Resolve the client ID in the system parameters At deployment time, Omnistrate exposes the application client ID through the [`$sys.deploymentCell.azure.managedWorkloadIdentities["managedIdentityIdentifier"].clientID`](https://docs.omnistrate.com/build-guides/system-parameters/#azure-specific-deployment-cell-parameters) system parameter. Replace `managedIdentityIdentifier` with the identity's `identifier` value. For the example above, use `$sys.deploymentCell.azure.managedWorkloadIdentities["blob-writer"].clientID`. The value is resolved for the target deployment cell, so the manifest does not need to contain a hard-coded subscription or client ID. #### Use the identity from a Kubernetes service account Annotate the service account used by the workload with the resolved client ID. For AKS Workload Identity, the `azure.workload.identity/client-id` annotation associates the Kubernetes service account with the Azure application: ``` apiVersion: v1 kind: ServiceAccount metadata: name: blob-writer namespace: blob-writer-ns annotations: azure.workload.identity/client-id: '{{ $sys.deploymentCell.azure.managedWorkloadIdentities["blob-writer"].clientID }}' ``` You can define such a service account in an amenity template or a plan spec (for use by a product instance). ### OCI Support for OCI managed workload identities is coming soon. # Monitoring with auto-recovery High availability is a critical component of your SaaS and it requires several measures to achieve high availability. In general, Omnistrate provides full support for your control plane, data plane (aka application) infrastructure and automated L1 support for your data plane failures. ## Control plane failures We will monitor, detect and recover any failures in your control plane to give you a 99.99% SLA. ## Data plane infrastructure failures Omnistrate automatically detects the following failures and seamlessly recovers from them: - Dead Process(es) - Machine failures - Network partitions - Degraded storage - Zonal failures If we notice these failures, we try basic recovery mechanisms, ex - restart the process or machine or replacing the machine. If we can't recover, we will alert your team with the configured mechanism to look into it further. ## Data plane non-infra issues In addition, Omnistrate provides mechanisms for you to detect and recover from process failures using healthcheck actionhook. In order to configure healthcheck actionhook, you can provide a check that we can use to validate the health of the process on a regular basis. As an example, let's say you want to verify liveness for your database application. You can perform a read after write query and make sure database is making progress. Note that you can specify different checks in the same healthcheck to make sure all components of your application are up and running. Let's say you also want to add a simple verification check to verify that your process is up and running, you can add a check using `ps` utility in addition to the above liveness check. If your process health check is failing, we will alert your team with relevant details to look into it. For more details on actionhooks, please see [here](https://docs.omnistrate.com/build-guides/actionhooks/index.md) For more details on integrating with Operational tools like PagerDuty. To learn more about integrations, please see [here](https://docs.omnistrate.com/build-guides/integrations/index.md) ## Understanding Healthy, Degraded, and Unhealthy The health badge in Omnistrate reflects more than just whether a pod is in `Running` state. - **Healthy**: Required platform checks, endpoint checks, and configured workload health signals are passing. - **Degraded**: The instance is partially available, or one of the required health signals is failing while some of the workload is still up. - **Unhealthy**: Required checks are failing and the instance is not considered operational. Common reasons an instance can appear `Degraded` or `Unhealthy` even when pods are running: - The customer-facing endpoint is not reachable yet - A Kubernetes `Service` or cloud load balancer is still pending - A custom healthcheck actionhook is failing - A dependent integration or sidecar is failing - The workload is running but not ready to serve traffic correctly When these states occur, use the workflow view, instance details, and [Debugging and Troubleshooting](https://docs.omnistrate.com/operate-guides/troubleshooting/index.md) together rather than relying on pod phase alone. ## Platform observability vs workload-specific monitoring stacks Omnistrate provides platform observability for deployment cells and fleet operations, but that is different from any workload-specific monitoring stack your application may require. If your chart or Operator expects components such as Prometheus Operator, Alertmanager, Grafana, or custom exporters to exist in the cluster, do not assume they are already present for your workload by default. Install and manage those cluster-level components at the deployment-cell layer by using [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md). If you want Omnistrate dashboards to include your workload metrics, use [Integrations](https://docs.omnistrate.com/build-guides/integrations/index.md) to define scrape targets, modeled metrics, and dashboards. # Operational automation Distributing your SaaS Product is only the first hurdle. Operating it at scale is even a bigger challenge as it requires large amount of engineering and support resources. With Omnistrate, you get: - Operational visibility - Inventory management - Monitoring - Alert management - Fleet operations - Tenant management ## Operational visibility To operate your SaaS, you need an Operational visibility into whats happening in your fleet from in-depth metrics, logging to alerting capabilities using your favorite tools. More specifically: - Comprehensive overview of your fleet with a dashboard that displays real-time alerts, health status, patching progress, and distribution across different versions for proactive management - For metrics and logging, please see [Integrations](https://docs.omnistrate.com/build-guides/integrations/#observability-integrations) - You can access Kubernetes Dashboard to check real-time operational status and health of Kubernetes clusters - For fleet-wide activity, you can see all the current operations happening in the fleet with details. ### Fleet dashboard The Fleet Dashboard provides comprehensive visibility into your SaaS fleet operations, offering real-time insights into deployment health, workflow status, patching progress, and operational metrics across all your deployment models and cloud environments. For more details, see [Fleet Dashboard](https://docs.omnistrate.com/operate-guides/fleet-dashboard/index.md). ## Inventory management You can search deployment instances, cells and explore your fleet knowing the exact software and hardware configuration provisioned for each of your deployments. ### Deployment instances Deployment instances are running instances of your SaaS product that have been provisioned for your customers. You can manage instances through multiple interfaces and perform various operations like stopping, restarting, scaling, and updating configurations. For more details, see [Deployment Instances](https://docs.omnistrate.com/operate-guides/deployment-instances/index.md). You can also adopt existing customer application installations and manage their upgrade path. For more information, see [Adopt Deployment Instances](https://docs.omnistrate.com/operate-guides/adopt-deployment-instances/index.md). ### Deployment cells Deployment cells are the foundational building blocks of Omnistrate's cellular architecture. These are isolated, self-contained units of compute infrastructure where your services run across multiple clouds and customer environments. For more details, see [Deployment Cells](https://docs.omnistrate.com/build-guides/deployment-cells/index.md). You can also adopt existing Kubernetes clusters as deployment cells. For more information, see [Adopt Deployment Cells](https://docs.omnistrate.com/operate-guides/adopt-deployment-cells/index.md). ### Deployment snapshots Omnistrate allows you to capture manual snapshots of your deployment instances to ensure data persistence and facilitate disaster recovery. These snapshots provide a point-in-time recovery head that remains independent of the source instance's lifecycle. For more details, see [Deployment Snapshots](https://docs.omnistrate.com/operate-guides/deployment-snapshots/index.md). ### BYOC cloud accounts Manage your customers' BYOC (Bring Your Own Cloud) accounts. BYOC addresses critical customer requirements around data sovereignty, security compliance, and cost control. For more details, see [BYOC Cloud Accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md). ## Automated operations Omnistrate automates common day-2 operations for your SaaS out of the box. ### Workflows Workflows represent automated sequences of operations that manage the lifecycle of your SaaS infrastructure. They orchestrate complex tasks such as provisioning, scaling, updating, and deprovisioning resources across your deployment infrastructure. For more details, see [Workflows](https://docs.omnistrate.com/operate-guides/workflows/index.md). ### Monitoring with auto-recovery Omnistrate provides comprehensive monitoring with auto-recovery capabilities for control plane, data plane infrastructure, and application health. The platform automatically detects and recovers from various types of failures. For more details, see [Monitoring](https://docs.omnistrate.com/operate-guides/monitoring/index.md). ## Fleet operations Fleet operations must be straightforward and efficient to ensure your team can respond quickly to incidents and routine maintenance tasks. Omnistrate centralizes operational tools, making it easy to execute actions such as patching, scaling, and rolling updates across your entire fleet with minimal effort. The platform's intuitive interfaces and automation capabilities reduce manual intervention, lower the risk of errors, and accelerate response times. ## Alert center Alert center allows you to configure notification types and channels to your needs. For detailed information on alert management and configuration, see [Alert Management](https://docs.omnistrate.com/operate-guides/alarms/index.md). ## Secure remote access Omnistrate provides secure remote access to any deployment cell (Kubernetes cluster) through the CLI, eliminating the need for complex networking configurations. This feature uses secure mTLS reverse tunneling and works with existing Kubernetes tools. For more details, see [Deployment Cell Access](https://docs.omnistrate.com/operate-guides/deployment-cell-access/index.md). # Debugging and Troubleshooting You can debug and troubleshoot deployments in Omnistrate by leveraging the tools and features described below. These capabilities provide you with detailed visibility into each step of the deployment process, enabling you to identify, diagnose, and resolve issues efficiently. ## Debugging Tools Overview Omnistrate provides debugging tools to help you diagnose and resolve deployment issues: | Tool | Purpose | When to use | | ------------------ | --------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | | **Debug Events** | View real-time workflow progress and errors | First step — identify which stage failed | | **Instance Debug** | Inspect the instance's resource dependency graph, workflow events, resource logs, metrics access, and resource-specific artifacts | Default starting point for Helm, Terraform, Compose, and Operator troubleshooting | ## Debug Events Omnistrate's **Debug Events** feature provides real-time, detailed insights into your deployments, allowing you to quickly pinpoint and resolve issues by offering a transparent view of each operation's progress and status. ### How Debug Events Works For each instance operation—such as create, modify, upgrade, start, stop, or delete—Omnistrate launches a dedicated workflow to carry out the task. This workflow moves through several defined stages, including bootstrap, storage, network, compute, deployment, and monitoring. Within each stage, actions are tracked as individual debug events in chronological order, giving a precise view of each action's execution status. This structure enables you to monitor the workflow's progression step-by-step and quickly locate the root cause of any issues. ### Debug Events Example Scenarios When a workflow encounters an issue, such as an invalid instance type parameter, debug events identify the specific error within the workflow. This detailed view enables you to diagnose and resolve issues swiftly, ensuring minimal disruption. By utilizing Omnistrate's Debug Events, you can streamline issue resolution, improve operational efficiency, and ensure a smoother service experience. ## Instance Debug For most day-to-day Helm, Terraform, Compose, and Operator troubleshooting, start with the instance debug view: ``` omnistrate-ctl instance debug ``` Use it to inspect the instance DAG, select the failing resource, and review resource-specific runtime details. The interactive view has a **Resource Details** tab for the deployment topology and a **Metrics** tab for Grafana dashboard metadata when metrics are enabled. This is the primary place to look for: - Workflow events for each resource - Application logs for Compose, Helm, Terraform, and Operator resources when log streams are available - Metrics dashboard access and published dashboard details when metrics are enabled - Rendered Helm values and Helm client logs - Terraform progress, rendered files, output, plan previews, execution logs, and operation history - Compose deployment API parameters and deployment output parameters - Operator deployment API parameters, deployment output parameters, and Operator CRD outputs For non-interactive automation, use JSON output: ``` omnistrate-ctl instance debug --output json ``` If you only need the metrics dashboard details, use: ``` omnistrate-ctl instance dashboard ``` ### Resource Debug Coverage | Resource type | What `instance debug` shows | | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | **Compose** | App logs, deployment API parameters, deployment output parameters, and workflow events | | **Helm** | Helm install/upgrade logs, app logs, rendered chart values, deployment API parameters, deployment output parameters, and workflow events | | **Terraform** | Progress, rendered Terraform files, Terraform output, live execution logs, app logs, operation history, plan previews, and workflow events | | **Operator** | App logs, deployment API parameters, deployment output parameters, Operator CRD outputs, and workflow events | Tip If logs or metrics are not shown, confirm that the relevant integration is enabled for the SaaS Product. Compose-based SaaS Products use `x-customer-integrations` or `x-internal-integrations`; Helm, Terraform, and Operator plans use `features.CUSTOMER` or `features.INTERNAL`. For configuration details, see [Integrations](https://docs.omnistrate.com/build-guides/integrations/index.md). ## Helm Troubleshooting Checklist For Helm-based SaaS Products, most repeated failures fall into a small set of categories. Start with `omnistrate-ctl instance debug ` and work through the following: 1. Confirm the rendered chart values match what you expect. 1. Read the Helm client logs for hook failures, timeouts, or Kubernetes API validation errors. 1. Review app logs for workload-level startup or readiness failures. 1. Check whether `Jobs`, hooks, `Services`, or load balancers are stuck in a pending state. 1. Look for leftover CRDs, finalizers, or namespaced resources if create or delete workflows keep failing. 1. Revisit Helm runtime flags such as `wait`, `waitForJobs`, `skipCRDs`, `upgradeCRDs`, and `timeoutNanos` if the chart behavior does not match the release lifecycle you need. For runtime flag details, see [Helm Charts Runtime Configuration](https://docs.omnistrate.com/build-guides/helm-charts-runtime-configuration/index.md). If you changed chart inputs or rendered artifacts, publish a new Plan version instead of repeatedly restarting the same workflow. See [Workflows](https://docs.omnistrate.com/operate-guides/workflows/#restarting-a-workflow-vs-publishing-a-new-plan-version). ## Compose Troubleshooting Checklist For Compose-based SaaS Products, use `omnistrate-ctl instance debug ` to open the Compose resource and work through the following: 1. Review app logs for container startup, health check, or dependency errors. 1. Confirm deployment API parameters resolved to the expected runtime values. 1. Confirm deployment output parameters contain the values other resources or customers expect. 1. Check workflow events to identify whether the failure happened during bootstrap, storage, network, compute, deployment, or monitoring. 1. Use the Metrics tab or `omnistrate-ctl instance dashboard ` to verify dashboard access when metrics are enabled. ## Operator Troubleshooting Checklist For Operator-based SaaS Products, use `omnistrate-ctl instance debug ` to open the Operator resource and work through the following: 1. Review app logs from the operator-managed workload. 1. Confirm deployment API parameters resolved to the expected custom resource inputs. 1. Inspect Operator CRD outputs and exported deployment output parameters. 1. Check workflow events to separate Omnistrate orchestration errors from operator reconciliation errors. 1. Use the Metrics tab or `omnistrate-ctl instance dashboard ` to verify dashboard access when metrics are enabled. ## Terraform Troubleshooting Checklist For Terraform-based SaaS Products, use `omnistrate-ctl instance debug ` to inspect rendered artifacts, `terraform plan` output, apply logs, provider configuration, and exported outputs. Work through the following: 1. Confirm rendered Terraform files contain the expected system parameters, API parameters, secrets, and dependency outputs. 1. Confirm the captured `terraform plan` matches the resources you expected to create, update, or delete. 1. Review apply logs for IAM permissions, quotas, unavailable regions or SKUs, networking inputs, and resource name conflicts. 1. Verify provider configuration and credentials, especially for BYOC, control-plane-targeted Terraform, and Nebius authentication. 1. Confirm Terraform outputs are available to downstream Helm, Operator, Kustomize, or Terraform resources. This visibility allows you to diagnose issues without needing direct access to the Terraform state or the execution environment. If a Terraform operation fails, the plan output helps you understand what Terraform attempted, and the apply logs show where it failed. Tip When debugging a Terraform failure, start by reviewing the plan output to confirm that the intended changes match your expectations. If the plan looks correct but the apply fails, check the apply logs for cloud provider errors such as quota limits, permission issues, or resource conflicts. Which surface to trust Use each surface for a different layer of the problem: - **Workflow view**: orchestration context, stage transitions, and the high-level error captured by the workflow - **Instance debug**: resource-level logs, metrics metadata, rendered artifacts, parameters, outputs, and operation history - **Live cluster or cloud state**: final confirmation of what actually landed The Workflow UI shows the status of the deployment workflow, while Instance Debug shows the current state of the deployed resources. These views may reflect different stages of the deployment lifecycle. For example, a workflow may time out after resources are successfully created, or a post-deployment validation workflow step may fail even though the resources are healthy. Use Instance Debug to investigate resource-level issues and Workflow Events to identify the workflow execution and step that produced the reported status. Omnistrate runs `terraform plan` and `terraform apply` automatically as part of the workflow. Instance Debug provides the rendered Terraform files and plan preview for review before you investigate apply logs. When Terraform source or inputs change, publish a new Plan version and trigger a fresh workflow; see [Restarting a Workflow vs Publishing a New Version](https://docs.omnistrate.com/build-guides/terraform-troubleshooting/#restarting-a-workflow-vs-publishing-a-new-version) for the exact behavior. For the full Terraform-specific workflow, see [Terraform Troubleshooting](https://docs.omnistrate.com/build-guides/terraform-troubleshooting/index.md). # Webhook Alarm Event Payloads This page documents the data Omnistrate sends to [webhook notification channels](https://docs.omnistrate.com/operate-guides/alarms/#webhook) for every alarm event. Use this reference when designing your webhook body template (`HTTP Request Body`). ### Template syntax `{{ $var. }}` resolves against the alarm event. Field names in `{{ $var. }}` are case-sensitive. ### Example template and rendered body HTTP Request Body ``` { "eventID": "{{ $var.id }}", "serviceID": "{{ $var.ServiceID }}", "eventName": "{{ $var.Name }}", "eventDescription": "{{ $var.Description }}", "eventType": "{{ $var.Type }}", "payload": "{{ $var.Payload }}", "timestamp": "{{ $var.CreatedAt }}" } ``` Rendered HTTP body (SuccessfulRestore example) ``` { "eventID": "event-gxocwiBnDH", "serviceID": "s-tVf1FKBQbl", "eventName": "Instance restored", "eventDescription": "SuccessfulRestore", "eventType": "SuccessfulRestore", "payload": { "instance_id": "instance-ep3zs6iwc", "Service": "My Database", "Organization": "Acme Corp", "Resource": "webserver", "ProductTier": "Enterprise", "subscription_id": "sub-abc123", "product_tier_id": "pt-xyz789", "source_instance_id": "instance-b7zds3mfg", "snapshot_id": "instance-ss-koet1ataz" }, "timestamp": "2026-04-20T17:35:07.043983Z" } ``` ## Available template variables The following fields are available for use in `HTTP Request Body` via `{{ $var. }}`. Association fields (`ServiceID`, `InstanceID`, etc.) are populated only when the alarm event relates to that entity. ### Identity and classification | Variable | Description | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------- | | `{{ $var.id }}` | Unique alarm event identifier | | `{{ $var.Type }}` | Specific event type (e.g. `SuccessfulRestore`) — see the per-category sections below | | `{{ $var.Category }}` | Category (`InstanceEvent`, `UserEvent`, `IdentityProviderEvent`, `BillingEvent`, `SystemEvent`, `DeploymentCellEvent`) | | `{{ $var.AlertType }}` | `Alarm` or `Notification` | | `{{ $var.Priority }}` | `Critical`, `High`, `Medium`, or `Low` | | `{{ $var.Name }}` | Short human-readable name | | `{{ $var.Description }}` | Human-readable description | ### Timing | Variable | Description | | ----------------------- | -------------------------------------------------- | | `{{ $var.CreatedAt }}` | ISO 8601 timestamp when the alarm event was raised | | `{{ $var.UpdatedAt }}` | ISO 8601 timestamp of the last update | | `{{ $var.ExpiryTime }}` | ISO 8601 timestamp when the alarm event expires | ### Associations | Variable | Description | | ----------------------------------- | -------------------------------------- | | `{{ $var.ServiceID }}` | ID of the associated SaaS product | | `{{ $var.ServiceEnvironmentID }}` | ID of the environment | | `{{ $var.ServiceEnvironmentType }}` | Environment type (`PROD`, `QA`, `DEV`) | | `{{ $var.InstanceID }}` | ID of the associated instance | | `{{ $var.ResourceID }}` | ID of the resource | | `{{ $var.ResourceVersion }}` | Version of the resource | ### Payload | Variable | Description | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `{{ $var.Payload }}` | The category-specific payload object. Render with `"{{ $var.Payload }}"` (quoted) to inline as raw JSON; resolves to `null` when absent. The fields available inside it depend on the event category — see the sections below. | ______________________________________________________________________ ## Instance events **Category:** `InstanceEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. These alarm events are delivered for instance lifecycle operations, health monitoring, scaling, backup/snapshot operations, and recovery actions. For the full list of event types in this category, see [Alarm Event Categories and Type](https://docs.omnistrate.com/operate-guides/alarms/#alarm-event-categories-and-type). ### Payload fields Most instance events deliver the following base fields inside `{{ $var.Payload }}`: | Field | Type | Description | | ----------------- | ------ | ------------------------------------- | | `instance_id` | string | ID of the affected instance | | `Service` | string | Name of the service | | `Organization` | string | Name of the end-customer organization | | `Resource` | string | Name of the resource | | `ProductTier` | string | Name of the plan | | `subscription_id` | string | ID of the subscription | | `product_tier_id` | string | ID of the plan | Example instance event payload ``` { "instance_id": "instance-12345678", "Service": "My Database", "Organization": "Acme Corp", "Resource": "PostgreSQL Primary", "ProductTier": "Enterprise", "subscription_id": "sub-abc123", "product_tier_id": "pt-xyz789" } ``` ### Monitoring event payload Health monitoring events (`UnhealthyInstance`, `UnhealthyIntegration`, `UnhealthyCustomerIntegration`) deliver a different set of fields inside `{{ $var.Payload }}`. These events expire after 1 hour rather than the 8-hour default used by other instance events. | Field | Type | Description | | -------------------------- | ------ | ------------------------------- | | `instance_id` | string | ID of the affected instance | | `instance_status` | string | Current status of the instance | | `service_id` | string | ID of the service | | `service_name` | string | Name of the service | | `service_environment_id` | string | ID of the service environment | | `service_environment_name` | string | Name of the service environment | | `product_tier_id` | string | ID of the plan | | `product_tier_name` | string | Name of the plan | | `cloud_provider_name` | string | Name of the cloud provider | | `region_code` | string | Cloud region code | Example UnhealthyInstance payload ``` { "instance_id": "instance-12345678", "instance_status": "UNHEALTHY", "service_id": "s-abc123", "service_name": "My Database", "service_environment_id": "se-def456", "service_environment_name": "Production", "product_tier_id": "pt-xyz789", "product_tier_name": "Enterprise", "cloud_provider_name": "aws", "region_code": "us-east-1" } ``` `HighCPUUsage` and `RecoveryStarted` events use the standard base payload and may include additional dynamic fields produced by the monitoring pipeline. ### Snapshot event payload Snapshot events (`FailedSnapshotCreate`, `SuccessfulSnapshotCreate`, `FailedSnapshotCopy`, `SuccessfulSnapshotCopy`, `FailedSnapshotDelete`, `SuccessfulSnapshotDelete`) deliver the base instance payload plus: | Field | Type | Description | | ------------- | ------ | ------------------ | | `snapshot_id` | string | ID of the snapshot | ### Restore event payload Restore events (`StartedRestore`, `FailedRestore`, `SuccessfulRestore`) deliver the base instance payload plus: | Field | Type | Description | | -------------------- | ------ | ---------------------------------------------------------------------------------- | | `source_instance_id` | string | ID of the original instance being restored from | | `snapshot_id` | string | ID of the snapshot used for the restore (present only for snapshot-based restores) | Example restore event payload ``` { "instance_id": "instance-12345678", "Service": "My Database", "Organization": "Acme Corp", "Resource": "PostgreSQL Primary", "ProductTier": "Enterprise", "subscription_id": "sub-abc123", "product_tier_id": "pt-xyz789", "source_instance_id": "instance-original-99", "snapshot_id": "snap-abc456" } ``` ______________________________________________________________________ ## User events **Category:** `UserEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. Delivered when end customers sign up, subscribe, get invited, are removed from a subscription, or delete their account in your SaaS product. ### Payload fields | Field | Type | Description | | -------------------- | ------- | ------------------------------------------------- | | `user_email` | string | Email address of the user | | `user_name` | string | Name of the user | | `user_id` | string | ID of the user | | `user_organization` | string | Name of the user's organization | | `is_user_enabled` | boolean | Whether the user account is enabled | | `service_name` | string | Name of the service (when applicable) | | `product_tier_name` | string | Name of the plan (when applicable) | | `subscription_id` | string | ID of the subscription (when applicable) | | `inviting_user_name` | string | Name of the inviting user (for invite events) | | `token` | string | Token associated with the event (when applicable) | | `invoice_id` | string | ID of the invoice (when applicable) | Example UserSignUp payload ``` { "user_email": "jane@acmecorp.com", "user_name": "Jane Smith", "user_id": "usr-abc123", "user_organization": "Acme Corp" } ``` Example UserSubscription payload ``` { "user_email": "jane@acmecorp.com", "user_name": "Jane Smith", "user_id": "usr-abc123", "user_organization": "Acme Corp", "is_user_enabled": true, "service_name": "My Database", "product_tier_name": "Enterprise", "subscription_id": "sub-xyz789" } ``` Example UserSubscriptionInvite payload ``` { "user_email": "bob@acmecorp.com", "user_name": "Bob Jones", "user_id": "usr-def456", "user_organization": "Acme Corp", "is_user_enabled": true, "service_name": "My Database", "product_tier_name": "Enterprise", "inviting_user_name": "Jane Smith", "subscription_id": "sub-xyz789" } ``` Example UserDeleted payload ``` { "user_email": "jane@acmecorp.com", "user_name": "Jane Smith", "user_id": "usr-abc123", "user_organization": "Acme Corp", "is_user_enabled": false } ``` ______________________________________________________________________ ## Identity provider events **Category:** `IdentityProviderEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. Delivered when an identity provider verification fails. ### Payload fields | Field | Type | Description | | ---------------------- | ------ | ------------------------------------------------ | | `IdentityProviderID` | string | ID of the identity provider | | `IdentityProviderName` | string | Name of the identity provider | | `IdentityProviderType` | string | Type of the identity provider | | `Status` | string | Current status of the identity provider | | `ClientID` | string | Client ID of the identity provider configuration | Example identity provider event payload ``` { "IdentityProviderID": "idp-abc123", "IdentityProviderName": "Okta SSO", "IdentityProviderType": "SAML", "Status": "FAILED", "ClientID": "0oa1b2c3d4e5f6g7h8" } ``` ______________________________________________________________________ ## Billing events **Category:** `BillingEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. Delivered for metering export operations and invoice generation. ### Payload fields When a billing event is associated with a specific service and plan, the payload includes: | Field | Type | Description | | ----------------- | ------ | ------------------- | | `ServiceName` | string | Name of the service | | `ServiceID` | string | ID of the service | | `Environment` | string | Environment type | | `ProductTierName` | string | Name of the plan | | `ProductTierID` | string | ID of the plan | Example billing event payload (with plan) ``` { "ServiceName": "My Database", "ServiceID": "s-abc123", "Environment": "PROD", "ProductTierName": "Enterprise", "ProductTierID": "pt-xyz789" } ``` For organization-level billing events (e.g., `InvoiceGenerateSuccess` without a specific plan), the payload contains only custom details provided at creation time. ______________________________________________________________________ ## System events **Category:** `SystemEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. Delivered for upgrade path operations. ### Payload fields The fields available inside `{{ $var.Payload }}` describe the upgrade path. Field names use PascalCase except for the two source-specific fields at the bottom. **Core fields (always present):** | Field | Type | Description | | ------------------- | ------ | ------------------------------------------------------ | | `ServiceID` | string | ID of the service | | `ProductTierID` | string | ID of the plan | | `UpgradePathID` | string | ID of the upgrade path | | `SourceVersion` | string | Source plan version | | `SourceVersionName` | string | Name of the source version | | `TargetVersion` | string | Target plan version | | `TargetVersionName` | string | Name of the target version | | `Status` | string | Current status of the upgrade path | | `Type` | string | Type of the upgrade path | | `CreatedAt` | string | Timestamp when the upgrade was created (ISO 8601) | | `ReleasedAt` | string | Timestamp when the upgrade was released (ISO 8601) | | `UpdatedAt` | string | Timestamp when the upgrade was last updated (ISO 8601) | **Progress fields:** | Field | Type | Description | | ----------------- | ------- | ------------------------------------------------------------ | | `TotalCount` | integer | Total instances eligible for upgrade | | `CompletedCount` | integer | Instances that completed the upgrade | | `FailedCount` | integer | Instances that failed the upgrade | | `PendingCount` | integer | Instances pending upgrade | | `SkippedCount` | integer | Instances that were skipped | | `InProgressCount` | integer | Instances currently upgrading | | `ScheduledCount` | integer | Instances scheduled for upgrade (present when a date is set) | **Optional fields (present when applicable):** | Field | Type | Description | | ----------------------- | ------- | ----------------------------------------------- | | `PlannedExecutionDate` | string | Planned execution date (ISO 8601) | | `CompletedAt` | string | Timestamp when the upgrade completed (ISO 8601) | | `CreatedBy` | string | Name of the user who created the upgrade | | `LastModifiedBy` | string | Name of the user who last modified the upgrade | | `LastRequestedAction` | string | Last maintenance action requested | | `NotifyCustomer` | boolean | Whether end customers are notified | | `MaxConcurrentUpgrades` | integer | Maximum concurrent upgrades | | `FailedInstanceReasons` | object | Map of instance IDs to failure details | | `StatusMessage` | string | Human-readable status message | **Event-specific fields:** | Field | Type | Description | | ---------------------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------- | | `scheduling_event_type` | string | `"scheduled"`, `"immediate"`, or `"reminder"` (for `UpgradeScheduled`) | | `completion_status` | string | `"SUCCESS"`, `"CANCELLED"`, or `"SKIPPED"` (for completion events) | | `pause_reason` | string | Machine-readable reason for the pause. Currently `"instances_failed"` (for `UpgradePaused`) | | `pause_reason_description` | string | Human-readable explanation of why the upgrade paused (for `UpgradePaused`) | | `failed_count_in_paused_batch` | integer | Number of instances that failed in the batch that triggered the pause (for `UpgradePaused`) | | `failed_instance_reasons` | object | Map of failed instance IDs to failure details for the batch that triggered the pause (for `UpgradePaused`) | | `failed_instance_reasons..reason` | string | Failure reason for the instance (for `UpgradePaused`) | | `paused_batch_number` | integer | Batch number that triggered the pause (for `UpgradePaused`) | | `paused_batch_size` | integer | Batch size used to calculate the pause failure ratio (for `UpgradePaused`) | | `paused_failure_ratio` | number | `failed_count_in_paused_batch / paused_batch_size` (for `UpgradePaused`) | | `remaining_instance_count` | integer | Number of instances still queued when the upgrade paused (for `UpgradePaused`) | Example UpgradeScheduled payload ``` { "ServiceID": "s-abc123", "ProductTierID": "pt-xyz789", "UpgradePathID": "up-abc123", "SourceVersion": "2.0", "SourceVersionName": "v2.0 Stable", "TargetVersion": "3.0", "TargetVersionName": "v3.0 GA", "Status": "Scheduled", "Type": "Major", "TotalCount": 15, "CompletedCount": 0, "FailedCount": 0, "PendingCount": 15, "SkippedCount": 0, "InProgressCount": 0, "ScheduledCount": 15, "CreatedAt": "2025-03-15T10:00:00Z", "ReleasedAt": "2025-03-15T10:00:00Z", "UpdatedAt": "2025-03-15T10:00:00Z", "PlannedExecutionDate": "2025-04-01T02:00:00Z", "CreatedBy": "admin@example.com", "NotifyCustomer": true, "scheduling_event_type": "scheduled" } ``` Example UpgradePaused payload ``` { "ServiceID": "s-abc123", "ProductTierID": "pt-xyz789", "UpgradePathID": "up-abc123", "SourceVersion": "2.0", "SourceVersionName": "v2.0 Stable", "TargetVersion": "3.0", "TargetVersionName": "v3.0 GA", "Status": "IN_PROGRESS", "Type": "Major", "TotalCount": 15, "CompletedCount": 6, "FailedCount": 1, "PendingCount": 8, "SkippedCount": 0, "InProgressCount": 0, "CreatedAt": "2025-03-15T10:00:00Z", "ReleasedAt": "2025-03-15T10:00:00Z", "UpdatedAt": "2025-03-20T02:15:00Z", "CreatedBy": "admin@example.com", "NotifyCustomer": true, "pause_reason": "instances_failed", "pause_reason_description": "Upgrade paused because 1 instance(s) failed in batch 3 while 8 instance(s) remained queued.", "failed_count_in_paused_batch": 1, "failed_instance_reasons": { "instance-abc123": { "reason": "Child workflow timeout (type: StartToClose)" } }, "paused_batch_number": 3, "paused_batch_size": 5, "paused_failure_ratio": 0.2, "remaining_instance_count": 8 } ``` ______________________________________________________________________ ## Deployment cell events **Category:** `DeploymentCellEvent` — Use `{{ $var.Payload }}` to receive the fields documented in this section. Delivered during the lifecycle of deployment cells. The payload structure varies by event type. ### Payload fields — bootstrap events The `DeploymentCellCreate*`, `DeploymentCellUpdate*`, and `DeploymentCellDelete*` event types deliver the following fields inside `{{ $var.Payload }}`: | Field | Type | Description | | ----------------- | ------ | ----------------------------- | | `host_cluster_id` | string | ID of the deployment cell | | `region` | string | Cloud region code | | `cloud_provider` | string | Cloud provider name | | `role` | string | Role of the deployment cell | | `org_id` | string | ID of the owning organization | | `root_user_id` | string | ID of the root user | Example DeploymentCellCreateStarted payload (bootstrap) ``` { "host_cluster_id": "hc-abc123", "region": "us-east-1", "cloud_provider": "aws", "role": "dataplane", "org_id": "org-xyz789", "root_user_id": "usr-abc123" } ``` ### Payload fields The `DeploymentCellStarted`, `DeploymentCellInProgress`, `DeploymentCellCompleted`, and `RepairingDeploymentCellStarted` event types deliver the following fields inside `{{ $var.Payload }}`: | Field | Type | Description | | --------------------- | ------ | ------------------------------------- | | `host_cluster_id` | string | ID of the deployment cell | | `host_cluster_status` | string | Current status of the deployment cell | | `host_cluster_role` | string | Role of the deployment cell | | `host_cluster_type` | string | Type of the deployment cell | | `organization_id` | string | ID of the owning organization | | `account_config_id` | string | ID of the cloud account configuration | On completion (`DeploymentCellCompleted`), these additional metrics fields are delivered: | Field | Type | Description | | ---------------------------- | ------ | ------------------------------------ | | `operation_duration_seconds` | number | Duration of the operation in seconds | | `operation_initial_status` | string | Status when the operation started | | `operation_final_status` | string | Status after the operation completed | Example DeploymentCellCompleted payload ``` { "host_cluster_id": "hc-abc123", "host_cluster_status": "READY", "host_cluster_role": "dataplane", "host_cluster_type": "kubernetes", "organization_id": "org-xyz789", "account_config_id": "ac-def456", "operation_duration_seconds": 842.5, "operation_initial_status": "PENDING", "operation_final_status": "READY" } ``` `HostClusterCleanup` events deliver the fields above plus the following cleanup-specific fields: | Field | Type | Description | | --------------------- | ------- | ----------------------------------------------------- | | `host_cluster_id` | string | ID of the deployment cell | | `host_cluster_status` | string | Current status of the deployment cell | | `host_cluster_role` | string | Role of the deployment cell | | `host_cluster_type` | string | Type of the deployment cell | | `organization_id` | string | ID of the owning organization | | `account_config_id` | string | ID of the cloud account configuration | | `root_user_id` | string | ID of the root user | | `region_id` | string | ID of the region | | `auto_cleanup` | boolean | Whether this was an automatic cleanup (`true`) | | `cleanup_reason` | string | Reason for cleanup (e.g., `inactive_deployment_cell`) | | `cleanup_details` | object | Nested details about the cleanup (see below) | The `cleanup_details` object contains: | Field | Type | Description | | ---------------------- | ------ | ------------------------------------------------- | | `inactivity_threshold` | string | Duration of inactivity that triggered cleanup | | `cleanup_time` | string | ISO 8601 timestamp when the cleanup was initiated | # Workflows ## What is a Workflow? A workflow in Omnistrate represents an automated sequence of operations that manage the lifecycle of your SaaS infrastructure and services. Workflows orchestrate complex tasks such as provisioning, scaling, updating, and deprovisioning resources across your deployment infrastructure. Workflows provide visibility into the execution of these operations, allowing you to monitor progress, troubleshoot issues, and ensure reliable service delivery to your customers. ## Workflow Operations Workflows can perform various types of operations: ### Workflow Type - **PROVISIONING**: Creates new instances and sets up required infrastructure - **DELETING**: Removes instances and cleans up associated resources - **SCALING**: Adjusts resource capacity up or down - **UPDATING**: Applies configuration changes or software updates - **START**: Initiates a stopped instance or service, making it active and available - **STOP**: Gracefully shuts down an instance or service, preserving state for later restart - **BACKUP**: Performs data backup operations to safeguard your information and enable point-in-time recovery - **RESTORE**: Recovers data from backups to restore a deployment to a previous state ### Workflow stages Each workflow consists of multiple stages, such as bootstrap, storage, network, compute, deployment, and monitoring. These stages are designed to run in parallel, allowing different parts of the workflow to execute simultaneously. This parallel execution improves efficiency and reduces the overall time required to complete the workflow, as independent tasks do not have to wait for each other to finish before starting. - **Bootstrap**: Initializes the foundational components of your service - **Storage**: Sets up persistent storage resources - **Network**: Configures networking components and connectivity - **Compute**: Provisions compute resources and containers - **Deployment**: Deploys your application components - **Monitoring**: Sets up observability and monitoring systems ### Control Operations Workflows support several control operations that can be performed through the Action dropdown: - **Pause**: Temporarily halt workflow execution - **Resume**: Continue a paused workflow - **Restart**: Re-execute a failed or terminated workflow from the beginning. This is useful when a transient issue caused the workflow to fail and you want to retry the entire operation. - **Terminate**: Stop and cancel a running workflow ## Workflow Visualization The workflow visualization provides a comprehensive view of each workflow activities with the following key sections: ### Status Overview The top section displays workflow status counters: - **In Progress**: Number of currently running workflows (shown with a gear icon) - **Succeeded**: Number of successfully completed workflows (shown with a checkmark icon) - **Failed**: Number of workflows that encountered errors (shown with a warning triangle icon) You can expand each step in the workflow visualization to view detailed information about its execution. This includes the current status, timestamps, and any error messages encountered during the operation. Expanding a step helps you diagnose issues quickly and understand the progress or failure points within the workflow. ### Troubleshooting Failed Workflows If a workflow fails, start with [Debug Events](https://docs.omnistrate.com/operate-guides/troubleshooting/#debug-events) to identify the failed stage. Then use [Instance Debug](https://docs.omnistrate.com/operate-guides/troubleshooting/#instance-debug) for resource-level logs, rendered artifacts, parameters, outputs, and operation history. ## Restarting a workflow vs publishing a new Plan version When a workflow starts, Omnistrate renders and captures the Plan, Terraform, or Helm inputs and other artifacts used by that operation. Restarting a failed workflow retries those captured artifacts; it does not normally re-read edits made afterward to the Plan specification, Terraform source, Helm inputs, or a pinned Git reference. Publish a new Plan version and trigger a fresh create, modify, or upgrade workflow if you changed any of the following: - The Plan specification - A `parameterDependencyMap` - The referenced Git branch, tag, or commit - Terraform or Helm artifacts that must be re-rendered If the referenced Git input is the moving head of a branch rather than a pinned commit or tag, a restart resolves the branch head again at restart time. For deterministic behavior, pin Git sources to a tag or commit SHA. When you need changed inputs to be rendered, publish a new Plan version and trigger a fresh create, modify, or upgrade workflow. # Governance Guides # API Keys ## Overview Omnistrate API keys are long-lived, org-scoped credentials that let you authenticate to the Omnistrate API without using a personal user account. They are designed for automation, CI/CD pipelines, service-to-service integrations, and powering the Customer Portal. Each API key: - Belongs to an **organization**, not to an individual user - Carries a fixed **role** (e.g. Admin, Service Editor) assigned at creation - Is returned in plaintext **exactly once** — at creation time - Is stored as an irreversible hash — the platform can never recover the plaintext - Has a recognizable `om_` prefix for easy identification in logs and configs ## Why Use API Keys Instead of User Accounts | Concern | User account | API key | | ------------------------ | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- | | **Credential lifecycle** | Tied to a person — breaks when employees leave or rotate passwords | Independent of any user — survives team changes | | **Least privilege** | Shares all permissions of the user | Scoped to a single role, fixed at creation (Root is never assignable) | | **Rotation** | Changing a password affects the user; only one credential per user | Multiple keys per org — create a replacement, then revoke the old one with zero downtime | | **Auditability** | Actions attributed to a person who may share credentials | Each key has its own identity in audit logs, with a unique ID and name | | **Secrets management** | Password must be treated as PII | API key is a machine secret — no email, no personal data | Tip Use API keys for any non-interactive workflow: CI/CD, automated deployments, portal backends, monitoring integrations, and scripts. Reserve user accounts for interactive console access. ## Creating an API Key ### Prerequisites - You must have **Root** or **Admin** role in your organization - The API Keys feature must be enabled for your organization ### Using the UI 1. In the Omnistrate console, navigate to **Settings → API Keys** 1. Click **Create API key** 1. Fill in the details: - **Name** (required) — a descriptive identifier, unique within the org (e.g. `ci-pipeline`, `saas-portal-prod`) - **Description** (optional) — what the key is used for - **Role** (required) — the highest role this key can assume. Root is never available; Admin is the maximum - **Expiry** (optional) — choose 30 days, 90 days, 1 year, or no expiry 1. Click **Create** 1. **Copy the key immediately** — it will not be shown again Warning The plaintext API key is displayed **only once** at creation time. Store it securely (e.g. in a secrets manager). If you lose it, you must create a new key and revoke the old one. ## Authenticating with an API Key API keys can be used to authenticate in three ways. All three produce the same session with the same permissions. ### Option 1: `X-API-Key` header (recommended for services) Send the API key in the `X-API-Key` header on every request. This is the simplest approach for service-to-service integrations and stateless scripts. ``` curl -X GET "https://api.omnistrate.cloud/2022-09-01-00/fleet/service" \ -H "X-API-Key: om_YOUR_API_KEY_HERE" \ -H "Content-Type: application/json" ``` **Example — list your services:** ``` export OMNISTRATE_API_KEY="om_YOUR_API_KEY_HERE" # List all services curl -s "https://api.omnistrate.cloud/2022-09-01-00/fleet/service" \ -H "X-API-Key: $OMNISTRATE_API_KEY" | jq . ``` **Example — describe a specific service:** ``` curl -s "https://api.omnistrate.cloud/2022-09-01-00/fleet/service/s-12345678" \ -H "X-API-Key: $OMNISTRATE_API_KEY" | jq . ``` ### Option 2: `Authorization: Bearer` header You can also pass the API key as a Bearer token. ``` curl -X GET "https://api.omnistrate.cloud/2022-09-01-00/fleet/service" \ -H "Authorization: Bearer om_YOUR_API_KEY_HERE" \ -H "Content-Type: application/json" ``` ### Option 3: Signin exchange (recommended for CLIs and long sessions) Exchange the API key for a short-lived JWT token via the signin endpoint. This way, the long-lived key crosses the wire only once per session. ``` # Exchange API key for a JWT TOKEN=$(curl -s -X POST "https://api.omnistrate.cloud/2022-09-01-00/signin" \ -H "Content-Type: application/json" \ -d '{ "email": "apikey@apikeys.invalid", "password": "om_YOUR_API_KEY_HERE" }' | jq -r '.jwtToken') # Use the JWT for subsequent requests curl -s "https://api.omnistrate.cloud/2022-09-01-00/fleet/service" \ -H "Authorization: Bearer $TOKEN" | jq . ``` Note When using the signin exchange, use the literal email `apikey@apikeys.invalid`. The email field is a routing hint — it is not validated as a real address. ## Using API Keys with the CTL ### Login with an API Key ``` # Recommended for CI/CD — pass via environment variable export OMNISTRATE_API_KEY="om_YOUR_API_KEY_HERE" omnistrate-ctl login # Pass via stdin (secure — key never appears in process list or shell history) echo "$OMNISTRATE_API_KEY" | omnistrate-ctl login --api-key-stdin ``` ### CI/CD with GitHub Actions Use the [`omnistrate-oss/setup-omnistrate-ctl`](https://github.com/omnistrate-oss/setup-omnistrate-ctl) action to install the CTL and authenticate with an API key in a single step. The action handles login, secret masking, and post-job cleanup (token revocation and credential removal) automatically. **Prerequisites:** store your API key as a GitHub Actions secret named `OMNISTRATE_API_KEY`. #### Basic usage ``` name: Deploy with Omnistrate on: push: branches: [main] jobs: deploy: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Setup Omnistrate CTL uses: omnistrate-oss/setup-omnistrate-ctl@v1 with: api-key: ${{ secrets.OMNISTRATE_API_KEY }} - name: Describe service run: omnistrate-ctl service describe --service-id ${{ vars.SERVICE_ID }} ``` The action passes the API key via stdin (never exposed in process lists or logs), masks it in all log output, and by default revokes the server-side refresh token when the job finishes. Tip Always prefer API keys over email/password for CI — they can be scoped, rotated, and revoked independently without affecting your user account. ## Managing API Keys ### List Keys In the Omnistrate console, navigate to **Settings → API Keys** to view all keys in your organization. ### Revoke a Key Revoking disables a key immediately. In the console, navigate to **Settings → API Keys**, find the key, and click **Revoke**. Revoking is idempotent — revoking an already-revoked key is a no-op. ### Delete a Key Delete permanently removes a revoked or expired key. Active keys cannot be deleted — revoke first. In the console, navigate to **Settings → API Keys**, find the revoked key, and click **Delete**. ### Rotate a Key There is no in-place regenerate. To rotate: 1. **Create** a new key with the same role 1. **Update** your integrations to use the new key 1. **Revoke** the old key This ensures zero-downtime rotation since both keys are valid during the transition. ## Security Best Practices - **Store keys in a secrets manager** — inject keys from a secrets manager for CI/CD, automation, and production. Never commit them to source control. - **Use the shortest practical expiry** — prefer 90 days for CI keys; avoid no-expiry when possible - **One key per integration** — name keys after their purpose (e.g. `github-actions-ci`, `portal-prod`) so you can revoke individually - **Prefer `X-API-Key` header** for stateless integrations and `signin exchange` for CLI sessions - **Revoke keys immediately** when an integration is decommissioned or a key may have been exposed - **Use the lowest required role** — most automation needs only Service Editor or Service Operator, not Admin # Audit Logs ## Overview Omnistrate Audit Logs provide an immutable record of all operations and activities performed within your environments. This centralized logging system maintains a complete audit trail for traceability and compliance purposes. The audit logs interface displays a chronological, immutable record of all operations across your SaaS offerings. These logs capture both system operations and customer actions, providing: - Complete traceability of all activities - Immutable audit trail for compliance requirements - Record of customer actions and system operations - Historical data for governance and security oversight ## Recorded Operations The audit logs capture all operations performed in your environment, including: - **Customer Actions** - All activities initiated by your customers (instance management, configuration changes, access events) - **System Operations** - Automated processes, maintenance tasks, and infrastructure events - **Administrative Actions** - Operations performed by service providers and administrators All logged operations are immutable and cannot be modified or deleted, ensuring data integrity for compliance and traceability purposes. ## Traceability and Compliance The immutable nature of audit logs ensures: - **Complete Traceability** - Every operation is recorded with timestamps, user information, and detailed context - **Compliance Support** - Immutable records meet regulatory requirements for audit trails - **Customer Activity Tracking** - Full visibility into customer actions and behaviors - **Data Integrity** - Logs cannot be altered, ensuring reliable historical records # Cloud Account Permission Requirements ## Overview In order to manage your customer deployments, Omnistrate assumes certain roles in your account for both BYOC and Hosted SaaS deployments. Specifically, Omnistrate separates the deployment process into two separate steps: - Service Bootstrap - Service Management ## Service Bootstrap Service bootstrap is the initial step Omnistrate takes to setup the platform to host your deployments. This is typically done only once per region for a given Cloud account. In the case of Hosted SaaS deployments where you host your customer deployments, the base platform is shared in that region and the configured cloud account. In the case of BYOC (Bring Your Own Cloud) deployments, the base platform is shared for all deployments of that particular customer / tenant who is bringing their account in that region. This step sets up the following: - AWS VPC / GCP Network / Azure Virtual Network - EKS / GKE / AKS Cluster - NGINX-backed NLB / GCP L4 LB - Your Service Deployment Agent ### AWS (bootstrap) #### Load Balancer Policy In order to deploy the NGINX-backed NLB, we bootstrap the [AWS Load Balancer Controller.](https://kubernetes-sigs.github.io/aws-load-balancer-controller/latest/) The permissions necessary for this controller are detailed [here](https://github.com/kubernetes-sigs/aws-load-balancer-controller/blob/main/docs/install/iam_policy.json). The policy includes: - **Service-linked Role Creation**: - Create service-linked roles for Elastic Load Balancing - **Read-only EC2 Permissions**: - Describe EC2 resources (VPCs, subnets, security groups, instances, etc.) - View network configurations and metadata - **Read-only ELB Permissions**: - Describe load balancers, target groups, listeners, and their attributes - **Security-related Permissions**: - Access to Cognito, ACM certificates, WAF, and Shield - Ability to view and manipulate security configurations - **Security Group Management**: - Create, modify, and delete security groups with specific tagging conditions - Manage security group ingress/egress rules - **Load Balancer Resource Management**: - Create and configure load balancers, target groups, and listeners - Register and deregister targets - Modify load balancer attributes and settings Most permissions include conditional statements based on resource tags (especially `elbv2.k8s.aws/cluster`) to limit actions to resources managed by the controller. This policy enables the AWS Load Balancer Controller to automatically provision and manage Application Load Balancers (ALB) and Network Load Balancers (NLB) for Kubernetes services. #### OmnistrateBootstrapPolicy The `OmnistrateBootstrapPolicy` is a custom policy that grants the necessary permissions for Omnistrate to bootstrap the platform. This policy is assumed through the Control Plane deployed in your AWS account in the case of BYOC deployments ensuring the trust relationship is established between the Control Plane in your account and the Data Plane in your customer's account. The full policy is available [here.](https://onboarding-cfv1.s3.us-west-2.amazonaws.com/account-config-setup-template.yaml) The policy includes: ##### IAM Role Management - **Role Creation with Boundary**: Can create, modify, and delete IAM roles, **but only when those roles use the specified permissions boundary** (`omnistrate-bootstrap-permissions-boundary`) - **Extensive IAM Administrative Actions**: Includes a wide range of IAM operations for read-only permissions for roles, policies, instance profiles, and OIDC providers within the AWS account ##### Service-Linked Roles - Can create specific service-linked roles for: - Amazon EKS - Amazon EKS Node Groups ##### EKS Management - Administrative access to Amazon EKS services (`eks:*`) ##### PassRole Permissions - Can pass roles to other AWS services in two circumstances: - Any role tagged with `omnistrate.com/managed-by: omnistrate` - Specific roles including: - Roles starting with `omnistrate-` - Custom NAT roles - EKS nodegroup service roles #### OmnistrateBootstrapPermissionsBoundary This IAM managed policy (`omnistrate-bootstrap-permissions-boundary`) defines the maximum permissions available to roles created during the Omnistrate bootstrapping process. It functions as a guardrail that limits what actions these roles can perform. The full policy is available [here.](https://onboarding-cfv1.s3.us-west-2.amazonaws.com/account-config-setup-template.yaml) The policy includes: - **IAM Management**: - Comprehensive IAM administrative capabilities, with a requirement that newly created roles must use this same permissions boundary - Full role, policy, instance profile, and OIDC provider management - **Core AWS Services (required)**: - Full access to EC2, EKS, Elastic Load Balancing, Auto Scaling, and CloudWatch - Limited KMS access (read-only) - **Storage Services (optional)**: - S3 permissions to manage workload buckets - Access to Amazon EFS resources - **Security & WAF (required for the Load Balancer Controller)**: - Full access to WAF, WAFv2, and Shield - Limited read access to Cognito, ACM, and IAM certificates - **Role Assumption & Passing**: - Can assume roles with `omnistrate-*` prefix - Can pass roles to services if they're tagged with `omnistrate.com/managed-by: omnistrate` - Can pass specific roles including EKS nodegroup roles This policy establishes a security boundary for the Omnistrate platform, allowing it to provision and manage infrastructure while providing guardrails on what actions it can perform within the AWS account. **All the policies setup in this phase can be easily removed once the platform is bootstrapped and the customer deployments are up and running. They only need to be restored in case of updates or new regions / accounts that need to be bootstrapped.** ## Service Management ### AWS (service management) #### OmnistrateInfrastructureProvisioningPolicy This IAM managed policy (`omnistrate-infrastructure-provisioning-policy`) defines permissions for the Service Deployment Agent to manage resources within the AWS account. This is necessary for deploying and managing customer workloads on the platform and is assumed ONLY by the Service Deployment Agent within the customer's account. The full policy is available [here.](https://onboarding-cfv1.s3.us-west-2.amazonaws.com/account-config-setup-template.yaml) The policy includes: - **IAM Read Access**: - Read-only access to all IAM resources in the account (`Get*` and `List*` actions) - **Specific Role Management**: - Permission to get, pass, and list policies for specific pre-defined roles: - EKS node group roles - Omnistrate EKS IAM roles - AWS service-linked roles for EKS - **IAM Resource Creation**: - Create and manage roles and policies within the `/omnistrate/` path - All role creation must use the `omnistrate-bootstrap-permissions-boundary` - Includes permissions to create, update, tag, and delete IAM resources - **Core Infrastructure Services**: - Full access to EC2, Elastic Load Balancing, EKS, and Auto Scaling - S3 access to workload buckets (optional) - Management of Amazon EFS resources (optional) - **Role Assumption**: - Can assume roles with the `omnistrate-*` prefix This policy enables Omnistrate's infrastructure provisioning component to create and manage AWS resources while operating within the constraints of the permissions boundary. It focuses on the core services needed for service deployment on Kubernetes while maintaining appropriate security guardrails. #### OmnistrateEC2NodeGroupIAMRole This IAM role is designed to support Amazon EKS deployments and is used by EC2 instances that serve as worker nodes in EKS clusters. The full policy is available [here.](https://onboarding-cfv1.s3.us-west-2.amazonaws.com/account-config-setup-template.yaml) **Key Characteristics**: - **Trust Relationship**: Can be assumed by the EC2 service - **Attached Policies**: - `AmazonEC2ContainerRegistryReadOnly`: Access to pull images from ECR - `AmazonEKSWorkerNodePolicy`: Core permissions for EKS worker nodes - `AmazonEKS_CNI_Policy`: Required for Kubernetes networking - `AutoScalingFullAccess`: Allows managing Auto Scaling groups - **Custom Inline Policy**: Allows describing EKS nodegroups - **Session Limit**: 1 hour (3600 seconds) - **Tagged**: `omnistrate.com/managed-by: omnistrate` #### OmnistrateEKSIAMRole This IAM role is designed to support Amazon EKS deployments and is used by the EKS control plane to manage AWS resources on behalf of the cluster. The full policy is available [here.](https://onboarding-cfv1.s3.us-west-2.amazonaws.com/account-config-setup-template.yaml) **Key Characteristics**: - **Trust Relationship**: Can be assumed by the EKS service - **Attached Policies**: - `AmazonEC2ContainerRegistryReadOnly`: Access to pull images from ECR - `AmazonEKSClusterPolicy`: Core permissions for EKS control plane - `AmazonEKSServicePolicy`: Legacy permissions (now included in cluster policy) - `AmazonEKSVPCResourceController`: Allows EKS to manage VPC resources - **Session Limit**: 1 hour (3600 seconds) - **Tagged**: `omnistrate.com/managed-by: omnistrate` Both roles are essential components for deploying and operating Amazon EKS clusters, with the first role supporting worker nodes and the second enabling the Kubernetes control plane to interact with AWS services. **The `OmnistrateInfrastructureProvisioningPolicy` can be removed once a customer deployment is up and running. It needs to be restored only in the case of any updates or new deployments that need to be managed.** # Omnistrate RBAC ## Overview Omnistrate RBAC allows your team members to assume some predefined roles. Note Please note that Omnistrate RBAC for your internal teams is completely different from your customer-facing RBAC for your customers. Omnistrate platform automatically builds RBAC for your customers when you build your SaaS Product using Omnistrate. Check the [Tenant Management](https://docs.omnistrate.com/tenant-management/overview/index.md) section for details on Customer RBAC. ## Organization Roles Here are the roles and associated permissions for different operations: | Organization Role | Build Side Service Access including (Pipeline, Build Service Definition, Product Tier) | Operate/Fleet Side Access | Account Config Access | User Invite Access (Access Control) | Billing Access | Service Plan Access | Resource Instance Access | Templates, Deployment Config, Image registry | | ----------------- | -------------------------------------------------------------------------------------- | ------------------------- | --------------------- | ----------------------------------- | -------------- | ------------------- | ------------------------ | -------------------------------------------- | | Root | CRUDL | CRUL | CRUDL | Invite/Uninvite except root | CRUDL | RL | CRUDL | CRUDL | | Admin | CRUDL | CRUL | RL | Invite/Uninvite except root | RL | RL | CRUDL | CRUDL | | Service Editor | CRUDL | RL | No Access | No Access | No Access | RL | CRUDL | CRUDL | | Service Operator | RL | CRUL | No Access | No Access | No Access | RL | CRUDL | RL | | Service Reader | RL | RL | No Access | No Access | No Access | RL | RL | RL | **Legend:** **C**: Create; **R**: Read/Describe; **U**: Update; **D**: Delete; **L**: List As an example, you may want to grant *Service Editor* role to your development team building control plane on top of Omnistrate and *Service Operator* to your platform teams to operate your SaaS using Omnistrate. ## Common assignment patterns - Use `Service Editor` for build and release workflows that update service definitions, Plans, pipelines, and registries. - Use `Service Operator` for day-2 operations on existing instances and deployment cells. - Use `Admin` or `Root` for cloud account onboarding or offboarding, organization access control, and other account-level configuration. ## Practical limitations - A user can hold only one organization role at a time. - If one automation flow needs both build or release privileges and cloud-account administration, use a higher-privilege role or split the workflow across separate bot users. - You can change a user's organization role directly from the Omnistrate console without needing to remove and re-invite the user. Navigate to **People** in the Omnistrate console, select the user, and update their role. ## Restrictions A given user can only be part of one organization. If a user is created without any invitation, it will have its own default organization. If a user is invited to an existing organization, that user will be part of that organization. If you would like to join a different organization, you need to be removed from your current organization and re-invited by the new organization you wish to join. Please be aware that during this transition, your original Omnistrate account will be deactivated, and you will need to create a new account. In other words, by moving to a new organization, you will lose access to any services or data associated with your original organization. If this process does not align with your needs, please contact [support@omnistrate.com](mailto:support@omnistrate.com) and we're here to help. ## SSO for the Omnistrate Console Omnistrate supports Single Sign-on (SSO) for the Omnistrate console, allowing your team members to authenticate using your organization's identity provider, such as Microsoft Entra ID, Okta, or any OpenID Connect-compatible provider. To configure SSO for your Omnistrate console, contact [support@omnistrate.com](mailto:support@omnistrate.com) with the following details: - Your Omnistrate organization ID - The identity provider you want to use (for example, Microsoft Entra ID) - Your OpenID Connect discovery URL or metadata endpoint Once configured, your team members can sign in to the Omnistrate console using their corporate credentials, enforcing your organization's authentication policies (including MFA) across all Omnistrate access. Note SSO for the Omnistrate console is separate from SSO for your Customer Portal. To configure SSO for your customers, see [Identity Providers (Single Sign-on)](https://docs.omnistrate.com/tenant-management/identity-providers/index.md). ## Disabling Username and Password Authentication After configuring SSO for your Omnistrate console, you can disable the default username/password authentication for your entire organization. This ensures that all team members must authenticate through your configured identity provider. To disable username/password authentication at the organization level, contact [support@omnistrate.com](mailto:support@omnistrate.com). Once enabled, this setting: - Prevents all organization members from signing in with username and password - Requires all users to authenticate through the configured identity provider - Applies to the Omnistrate console only (Customer Portal authentication is configured separately) Warning Before disabling username/password authentication, ensure that at least one identity provider is configured and tested for your organization. Otherwise, your team members will be locked out of the Omnistrate console. # Security and Compliance at Omnistrate Omnistrate delivers enterprise-grade security and compliance for your SaaS products across all deployment channels. Our platform is designed to help you implement robust security controls, maintain governance, and meet industry standards with ease. ## Our Security Practices - **SOC2 Type II Compliance**: Omnistrate's control plane is SOC2 Type II compliant. We have documented and implemented all required controls to help you accelerate compliance for your SaaS product. For more details, see our [Type I](https://blog.omnistrate.com/posts/26) and [Type II](https://blog.omnistrate.com/posts/33) press releases. - **Penetration Testing**: We conduct regular penetration tests and make reports available upon request to ensure ongoing security and risk mitigation. - **Vulnerability Reporting**: If you discover a security issue or vulnerability, please contact us at [security@omnistrate.com](mailto:security@omnistrate.com). We will coordinate with you to securely address and resolve any concerns. ## Key Security Features - [Cloud Account Permissions](https://docs.omnistrate.com/governance-guides/cloud-account-permissions/index.md): Permissions required to access your and your customers' cloud accounts, following minimum privilege principles - [Role-Based Access Control (RBAC)](https://docs.omnistrate.com/governance-guides/omnistrate-rbac/index.md): Enforce security policies to restrict access to resources and actions based on user roles and organizational context - [API Keys](https://docs.omnistrate.com/governance-guides/api-keys/index.md): Long-lived, org-scoped credentials for automation, CI/CD, and service-to-service authentication - [Secrets Management](https://docs.omnistrate.com/dev-ops-guides/secrets/index.md): Secure handling of sensitive configuration - [Customer Networks](https://docs.omnistrate.com/runtime-guides/customer-networks/index.md): Network isolation and VPC integration - [SSO / Identity Providers](https://docs.omnistrate.com/tenant-management/identity-providers/index.md): Enterprise authentication and identity management - [Operational Status Page](https://docs.omnistrate.com/build-guides/integrations/#operational-status): Communicate real-time status to your users - [Audit Logs](https://docs.omnistrate.com/governance-guides/audit-logs/index.md): Track user actions and system events across your SaaS products for accountability and compliance ## Compliance Resources The resources in this section are designed to help you create and certify your own self-service portal, ensuring it meets industry security and compliance standards. - **SOC2 for your control plane**: Accelerate SOC2 compliance for your SaaS product using Omnistrate's certified control plane and documented controls. - **Security questionnaire**: Access a common compliance questionnaire report, including resources like the [AWS FTR checklist](https://aws.amazon.com/partners/foundational-technical-review/), to streamline your review process. - **Pen test report**: Review details about our penetration testing program. If you require a custom report, please contact [support@omnistrate.com](mailto:support@omnistrate.com) for assistance. ## Reporting Security Issues Note If you have a security concern or believe you have found a vulnerability in any part of our infrastructure, please contact us at [security@omnistrate.com](mailto:security@omnistrate.com). We will work with you to coordinate the secure exchange of sensitive information. # FinOps Guides # End-to-End Billing ## End-to-End Billing with Omnistrate Omnistrate provides a complete solution for usage-based billing for your SaaS Product. It provides the following features: - Metering: Collect usage data for your SaaS Product - Aggregation: Aggregate usage data across all the nodes of an instance and per customer - Invoicing: Generate invoices for your customers - Payment: Integrate with payment processors to collect payments - Pricing plans: SaaS Providers can configure pricing plans for their SaaS Product and enforce quotas/restrictions per service-plan. - Notifications: Notify customers about their invoices and show them the current usage. - Payment collection: Receive payments to your account when customers accept their invoices. Billing can be enabled for your Plan with the following steps: 1. Enable Tenant Billing 1. Connect Omnistrate with your Stripe Account (or bring your own billing provider) 1. Configure pricing and billing provider for your Plan 1. Configure quotas for your Plan (optional) Invoices, Notifications and Payment collection are automated. By enabling this, your customers will be able to track their usage and invoices through the customer portal. Your customers will have the ability to configure payments with Stripe and pay their invoices through the Customer Portal. Note If a customer user is suspended, they can still sign in to the Customer Portal to view billing information and perform payment-related activities, including resolving outstanding invoices. Suspension blocks operational activities such as creating, updating, or managing deployment instances. ## Enable Tenant Billing At Account Level To enable tenant billing at account level, please navigate to the *"FinOps Center > Tenant Billing"* section. As a SaaS Provider, you can request Omnistrate to enable the feature. When the feature is enabled, you will see the *"Add Billing Provider"* button. Click on it to configure your billing provider. You can choose between using Omnistrate's billing provider(Stripe) or bringing your own billing provider. ### OmniBilling (Stripe) Omnistrate provides a built-in billing provider that integrates with Stripe. This allows you to manage your customer billing workflow through Omnistrate, including usage metering, pricing configuration, quota management, automatic invoice generation and notifications, invoice management, and integration with payment processors. You can configure your Stripe account and set up billing for your SaaS Product. To complete the Stripe enablement you need to configure your Stripe Account using “Stripe Connect”. If you don't have a Stripe Account, you must set up a standard Stripe account. Enter all required business information, including address and payment details, so the platform can process payments successfully. For more information on how to create a Stripe account you can review the [Stripe's Getting Started](https://docs.stripe.com/get-started/account#get-started-with-a-stripe-account) guideline. Warning Make sure to configure and save the [Customer Portal](https://dashboard.stripe.com/settings/billing/portal) properties, including your logo, data that will be requested from you customers (for instance Billing address for tax purposes) and accepted payment methods. For more details on how we establish the connection, see the [Stripe Documentation](https://docs.stripe.com/connect/standard-accounts) on how Stripe Connect works. ### BYO-Billing Provider (Bring Your Own Billing Provider) If you have an existing billing provider, you can integrate it with Omnistrate. This allows you to collect usage metering data from Omnistrate and integrate them with your billing system. You can enable the ["Omnistrate Metering"](https://docs.omnistrate.com/fin-ops-guides/metering/index.md) feature at the plan level to collect usage data and export it to your S3/GCS bucket, which can then be processed by your billing system. This allows you to use Omnistrate for usage metering while keeping your existing billing system intact. You will need to configure below properties to better display your billing provider in the Customer Portal: - **Name**: The name of the billing provider. - **Logo**: The logo of the billing provider. - **Balance Due Link**: The link to the balance due page of the billing provider. This is where your customers will be redirected to pay their invoices. ### What Omnistrate automates vs what remains in your billing system When you use a BYO billing provider or marketplace integration, Omnistrate handles the usage collection side of the workflow and exposes billing context in the platform. Your external billing system remains the system of record for invoice generation or marketplace submission. Omnistrate handles: - Plan-level pricing and billing-provider configuration - Usage metering collection - Exporting normalized usage records to your S3 or GCS bucket - Surfacing customer, subscription, environment, and plan metadata with each usage record - Presenting your provider name, logo, and balance-due link in the Customer Portal Your billing or marketplace integration handles: - Transforming exported records into the target billing-provider schema - Submitting usage to a marketplace or creating invoices in your own billing system - Payment reconciliation, taxes, disputes, and any downstream finance workflows outside Omnistrate ### Typical external billing or marketplace flow For external billing systems and cloud marketplace adapters, the common pattern is: 1. Configure a BYO billing provider in Omnistrate. 1. Enable metering export on the relevant Plans. 1. Optionally pass an `externalPayerId` when creating instances so exported usage can be correlated with your external payer record. 1. Consume the S3 or GCS exports in your billing or marketplace adapter. 1. Transform and submit usage to the downstream billing system. 1. Use the billing provider metadata in the Customer Portal so customers know where balances and invoices are managed. For a reference implementation of this pattern for Clazar-backed cloud marketplace billing, see the [usage-export-clazar-recipe](https://github.com/omnistrate-community/usage-export-clazar-recipe). For more on marketplace-oriented flows, see [Cloud Marketplace Integrations](https://docs.omnistrate.com/fin-ops-guides/marketplace/index.md). For record format and export layout, see [Measure your SaaS Product usage](https://docs.omnistrate.com/fin-ops-guides/metering/index.md). ## Configure Pricing Once the billing is enabled at an account level, you can configure the pricing for your SaaS Product at a per-plan basis. You can configure the pricing for the following pre-set dimensions: - CPU Cores - Allocated Memory - Allocated Storage (depending on the service definition) - GPU - Replicas - Deployment Cells In addition, you can restrict your customers from creating an instance if they had not configured payment. You can configure the pricing for your Plan through the UI or through the compose spec. - Using the UI, navigate to the *"Build Service > Plans"* section and modify the Plan you want to configure. - Using the compose spec, provide the `pricing` specifications under the `x-omnistrate-service-plan` section in your [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md). - Using the plan spec, provide the `pricing` specifications in your [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). Here is an example of how to configure the pricing for your Plan: ``` pricing: - dimension: cpu unit: cores timeUnit: hour price: 0.01 - dimension: memory unit: GiB timeUnit: hour price: 0.05 - dimension: storage unit: GiB timeUnit: hour price: 0.02 - dimension: gpu unit: millicores timeUnit: hour price: 0.04 - dimension: replica timeUnit: hour price: 0.10 - dimension: deploymentCell timeUnit: hour price: 0.10 ``` - `pricing`: The usage-based pricing model of the Plan. It will be used to generate invoices at the end of the month. Make sure to enable tenant billing first before using this feature. For a detailed guide, see [here](https://docs.omnistrate.com/fin-ops-guides/billing/index.md). Please specify the pricing for the following dimensions: - `cpu`: The pricing is based on the CPU usage. - `unit`: (Optional) For now we only support `cores`. Defaults to `cores` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. - `memory`: The pricing is based on the amount of memory (RAM) used. - `unit`: (Optional) For now we only support `GiB`. Defaults to `GiB` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. - `storage`: The pricing is based on the amount of allocated storage. - `unit`: (Optional) For now we only support `GiB`. Defaults to `GiB` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. - `gpu`: The pricing is based on allocated GPU usage. - `unit`: (Optional) For now we only support `millicores`. Defaults to `millicores` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. - `replica`: The pricing is based on the Replica usage. - `unit`: (Optional) For now we only support `replicas`. Defaults to `replicas` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. - `deploymentCell`: The pricing is based on Deployment Cell usage. - `unit`: (Optional) For now we only support `deploymentCells`. Defaults to `deploymentCells` if omitted. - `timeUnit`: (Optional) For now we only support `hour`. Defaults to `hour` if omitted. When an instance is stopped or stopping, allocated persistent storage can remain billable because the storage allocation continues to exist. Compute-like dimensions such as CPU, memory, GPU, replica, and deployment cell usage are not billed while the instance is stopped. Info To enable replica-based billing for custom tenancy plans, you need to explicitly label the pods you want to include in customer billing. For more details, please see the [section below](#enable-replica-based-billing-in-custom-tenancy-plans). ## Configure Billing Providers You will also need to choose the billing providers for your Plan. You can choose any billing providers that you have configured at the account level. And set the default billing provider that will be used for your new subscribers. - Using the UI, navigate to the *"Build Service > Plans"* section and modify the Plan you want to configure. - Using the compose spec, provide the `billingProviders` specifications under the `x-omnistrate-service-plan` section in your [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md). - Using the plan spec, provide the `billingProviders` specifications in your [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). Here is an example of how to configure the billing providers for your Plan: ``` billingProviders: - name: stripe externalProductID: "prod_123" enablePaywall: false disableInvoiceGeneration: false isDefault: true - name: BYO Billing Provider # The name you set for your BYO billing provider when configuring tenant billing at account level isDefault: false ``` - `billingProviders`: Configure billing providers for your service plan. Each provider supports different payment methods and billing integrations. All specified providers must be enabled for tenant billing at the account level. - `name`: The billing provider name. Use `stripe` (case-insensitive) for Stripe integration or the name you set for your bring-your-own (BYO) billing provider. - `isDefault`: (Optional) Whether this provider is the default. Only one provider can be set as default. Defaults to `false`. - `externalProductID`: (Stripe only) You can specify the billing product ID for your Plan. This can be used to identify the Plan in your billing system. The billing product ID is a unique identifier for the Plan and is used to track usage and generate invoices. - `enablePaywall`: (Stripe only) Whether a valid payment method is required for your customers to create resource instances. Possible options are: - `true`: Your customers must provide a valid payment method before consuming the service. - `false`(default): Your customers can consume the service without providing a valid payment method. - `disableInvoiceGeneration`: (Stripe only) Whether Omnistrate should skip Stripe invoice generation for this Plan. When set to `true`, Omnistrate continues to collect usage but does not create Stripe invoices for usage on this Plan. Note Above configurations that applied to the Plan level won't be applied to the existing subscriptions. You can always manage and update the existing subscriptions to apply the new configurations at the `FinOps Center > Tenant Pricing` section. ## Configure Quotas As part of the billing configuration, you can set quotas for your customers. If you use Stripe as your billing provider, you can restrict your customers from creating an instance if they had not configured payment by enabling the `enablePaywall` option above. You can also set a maximum number of instances that your customers can create. This can be useful if you want to limit the number of instances that your customers can create based on their subscription plan. - Using the UI, navigate to the *"Build Service > Plans"* section and modify the Plan you want to configure. - Using the compose spec, provide the `maxNumberOfInstancesAllowed` specifications under the `x-omnistrate-service-plan` section in your [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md). - Using the plan spec, provide the `maxNumberOfInstancesAllowed` specifications in your [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). Here is an example of how to configure the maximum number of instances for your Plan: ``` maxNumberOfInstancesAllowed: 5 ``` Note Above configurations that applied to the Plan level won't be applied to the existing subscriptions. You can always manage and update the existing subscriptions to apply the new configurations at the `FinOps Center > Tenant Pricing` section. ## Configure Invoices Once the billing is enabled, Omnistrate will automatically generate invoices in Stripe for your customers based on their usage at the end of the month, unless Stripe invoice generation is disabled for the Plan. The invoices will be generated based on the usage data collected by Omnistrate and will be created in Draft for your review. Disabling invoice generation does not delete or cancel existing Stripe invoices. Optionally, you can provide a billing product ID for your Plan, which will be used to create the Plan product in Stripe. - Using the UI, navigate to the *"Build Service > Plans -> Modify Plan"* and configure the `External product ID` in the `Billing` section. - Using the compose spec, find the `externalProductID` specifications in the [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md) under the `x-omnistrate-service-plan` section. - Using the plan spec, find the `externalProductID` specifications in the [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/#billing-provider-schema). To manage the monthly invoices and approve them, navigate to "Manage Fleet" section of the UI and click on "Manage Invoices". If you would like to auto-approve all your generated invoices after a certain period, please request through support. Note Stripe will allow you to modify invoices in Draft mode. After the invoices are sent to customers you can create a revision, invalidating the previous invoice and creating a new one. ## Customer Notifications Stripe will automatically notify your customers about their Open invoices. The notifications will be sent to the email address provided by your customers. Your customers will be able to track their usage and invoices through the customer portal. Your customers will have the ability to configure payments with Stripe and pay their invoices through the Customer Portal. ## Tax Handling When using Stripe as your billing provider, you can configure tax collection for your customers. Omnistrate integrates with Stripe's tax capabilities to support automatic tax calculation and collection on invoices. ### Configuring taxes in Stripe To enable tax handling for your SaaS Product: 1. **Enable Stripe Tax**: Navigate to the [Stripe Tax settings](https://dashboard.stripe.com/settings/tax) in your Stripe dashboard and enable automatic tax collection 1. **Configure your tax registrations**: Add your tax registration numbers for the jurisdictions where you are required to collect tax 1. **Set your product tax codes**: In your Stripe dashboard, assign the appropriate [tax codes](https://stripe.com/docs/tax/tax-codes) to your products to ensure correct tax rates are applied ### Customer location information To calculate the correct tax rates, Stripe Tax requires sufficient customer location information (billing address recommended) from your customers. In the [Customer Portal settings](https://dashboard.stripe.com/settings/billing/portal) in Stripe, enable address collection (for example, set **Address collection** to *Required* and choose *Billing* as the address type) so that Stripe can collect this information. Warning Make sure the Customer Portal in Stripe is configured to collect customer address information (for example, via the **Address collection** setting). Without sufficient customer location information, such as a billing address, Stripe may not be able to calculate applicable taxes. ### How tax appears on invoices When tax is configured, Omnistrate-generated invoices in Stripe automatically include: - Line items for resource usage (CPU, memory, storage, replicas) - Applicable tax amounts calculated by Stripe based on the customer's location information and your tax registrations - Tax breakdown showing the tax rate and jurisdiction Customers see the tax details on their invoices in the Customer Portal. The total amount due includes both usage charges and applicable taxes. ### Tax-inclusive vs. tax-exclusive pricing In Stripe, tax behavior (whether prices are tax-inclusive or tax-exclusive) is configured on your **Stripe Prices** for each product or plan. You can review and manage your overall tax configuration in the [Stripe Tax settings](https://dashboard.stripe.com/settings/tax), but inclusive vs. exclusive behavior is typically set per Price. - **Tax-exclusive** (default): Tax is added on top of the Stripe Prices you configure for your Plan pricing. Customers see the base price plus tax. - **Tax-inclusive**: The Stripe Prices you configure already include tax, and Stripe calculates the tax amount from that inclusive price. Refer to the [Stripe Tax behavior documentation](https://stripe.com/docs/tax/products-prices-tax-categories-tax-behavior#tax-behavior) for details on how tax behavior is applied per Price. Tip Review the [Stripe Tax documentation](https://stripe.com/docs/tax) for detailed guidance on tax compliance, supported countries, and tax registration requirements. ## Enable replica-based billing in custom tenancy plans This configuration is **optional** and only required if you plan to charge your customers based on the number of replicas in your pricing model. If replica-based billing is not part of your pricing strategy, you can skip this step entirely. For **OMNISTRATE_DEDICATED_TENANCY** and **OMNISTRATE_MULTI_TENANCY** plans, all pods are automatically included in customer billing. For **CUSTOM_TENANCY** plans (e.g., Helm, Operator, Kustomize), you must explicitly label the pods you want to include in customer billing. Add the following **pod label** (not an annotation) to your deployment manifests: ``` omnistrate.com/include-customer-billing: "true" ``` Only pods with this label will be billed. Pods without it will be excluded from customer billing. ### Example: Adding the Label in a Deployment Manifest ``` apiVersion: apps/v1 kind: Deployment metadata: name: example-app spec: replicas: 3 selector: matchLabels: app: example-app template: metadata: labels: app: example-app omnistrate.com/include-customer-billing: "true" spec: containers: - name: example-container image: nginx:latest ``` In this example, the pod template includes the `omnistrate.com/include-customer-billing: "true"` label under `metadata.labels`. This ensures all pods created by this deployment are counted for billing. # Cost Control & Insights ## Overview Managing infrastructure costs efficiently is critical for any SaaS business. Cost Control & Insights provides complete visibility and control over your infrastructure and tenant costs across **AWS, Azure, GCP, and OCI**, allowing you to analyze, optimize, and reduce expenses while maintaining performance and scalability. ## Multi-Cloud Cost Tracking Omnistrate tracks infrastructure costs across all supported cloud providers: | Cloud Provider | Cost Tracking | Granularity | | -------------- | ------------- | --------------------------------------------- | | AWS | Supported | Per-environment, per-deployment, per-resource | | Azure | Supported | Per-environment, per-deployment, per-resource | | GCP | Supported | Per-environment, per-deployment, per-resource | | OCI | Supported | Per-environment, per-deployment, per-resource | Cost data is collected automatically from each cloud provider and aggregated into a unified view, regardless of whether you deploy in your own cloud (BYOC) or use Omnistrate-hosted infrastructure. ## Accessing Cost Insights You can access cost insights from the Omnistrate console: 1. Navigate to **FinOps** > **Cost Insights** in the left sidebar. 1. Select the time range you want to analyze. 1. Filter by environment, deployment, or resource to drill down into specific costs. ## Cost Dashboard The Cost Dashboard provides a visual, real-time overview of your infrastructure spending. Access the dashboard from **FinOps** > **Cost Insights** in the Omnistrate console. ### Dashboard views The Cost Dashboard includes the following views: - **Overview**: A high-level summary of total infrastructure costs across all cloud providers, with trend indicators showing month-over-month changes. - **By cloud provider**: A breakdown of costs per cloud provider (AWS, Azure, GCP, OCI), allowing you to compare spending across clouds. - **By environment**: Compare costs across your environments (development, staging, production) to identify opportunities for optimization. - **By tenant**: Attribute costs to individual customers to understand per-tenant margins and identify your most resource-intensive workloads. ### Time range and filtering - Select a **time range** (last 7 days, last 30 days, last 90 days, or a custom range) to analyze spending trends. - **Filter** by environment, deployment, resource type, or cloud provider to drill down into specific cost areas. - **Export** cost data for further analysis or integration with external FinOps tools. ### Cost trends and alerts The dashboard surfaces cost trends automatically: - **Spending anomalies**: Unexpected cost spikes are highlighted so you can investigate and take action. - **Growth trends**: Track how costs are changing over time to forecast future spending. - **Budget tracking**: Monitor spending against budgets to stay within financial targets. ## Cost Attribution Use **infrastructure tags** to categorize and analyze spending across projects, teams, and tenants: - **Per-environment costs**: Compare spending across development, staging, and production environments. - **Per-deployment costs**: Track costs for individual deployments to identify the most expensive workloads. - **Per-resource costs**: Break down costs by compute, storage, and networking for each resource. - **Per-tenant costs**: Attribute infrastructure costs to individual customers for accurate cost allocation and margin analysis. ## Optimization Strategies Omnistrate provides several built-in mechanisms to reduce infrastructure costs: ### Spot and reserved instances Leverage discounted compute pricing from cloud providers: - **Spot instances**: Use spare cloud capacity at significantly reduced rates for fault-tolerant workloads. - **Reserved instances**: Commit to longer-term usage for predictable workloads and receive discounted pricing. ### Auto-scaling and auto-stop Dynamically adjust resources based on real-time demand: - **Auto-scaling**: Automatically scale resources up or down based on CPU, memory, or custom metrics using [autoscaling](https://docs.omnistrate.com/runtime-guides/autoscaling/index.md). - **Auto-stop**: Automatically stop idle deployments to eliminate costs from unused resources. ### Multi-tenancy Maximize infrastructure efficiency through built-in multi-tenancy support: - Share underlying infrastructure across multiple tenants to reduce per-customer costs. - Use Omnistrate's tenancy models to balance isolation requirements with cost efficiency. ### Deploy in your cloud Deploy in your own cloud account or your customer's cloud account to: - Use existing cloud credits and committed-use discounts. - Avoid data transfer costs between providers. - Take advantage of negotiated enterprise pricing. # Send Custom Usage Metrics As a SaaS Provider, use the custom usage metering API to report application-specific usage for a customer subscription. Identify the subscription directly with `subscription-id`, indirectly with a linked Resource `instance-id`, or with both identifiers. First obtain a temporary metrics endpoint and token from Omnistrate, then send your usage events to that endpoint. ## How the Integration Works 1. Call the Omnistrate organization API with your API key or token. 1. Receive a complete metrics endpoint URL, a temporary token, and its expiration time. 1. Send a JSON array of custom usage events to the returned endpoint with the temporary token. Warning Keep your Omnistrate credential and the returned metrics token secret. Do not place credentials in source control, logs, screenshots, or client-side application code. Request a new metrics token before the current token reaches its `expiresAt` time. ## Prerequisites Before sending custom usage metrics, you need: - An Omnistrate API key or token authorized to retrieve custom metrics credentials. Root, Admin, and Service Operator roles are supported. - The customer subscription that should receive the usage. - For instance-based reporting, a Resource instance associated with that subscription. - A custom metric definition configured for the subscription's product tier. - A unique `idempotency-id` for every logical usage event. - An RFC 3339 timestamp and one or more integer metrics for every event. - `curl` for the examples below. The optional shell automation example also uses `jq`. ## Get the Metrics Endpoint and Token Call the organization API with an Omnistrate API key or token. The credential must have the Root, Admin, or Service Operator role. The optional `expiresInSeconds` query parameter sets the lifetime of the metrics token. Its value must be between 1 and 86,400 seconds (24 hours). If omitted, the token is valid for 24 hours. ``` GET https://api.omnistrate.cloud/2022-09-01-00/custom-metrics/endpoint ``` ``` export OMNISTRATE_API_KEY_OR_TOKEN="" curl --silent --show-error --request GET \ "https://api.omnistrate.cloud/2022-09-01-00/custom-metrics/endpoint?expiresInSeconds=86400" \ --header "Authorization: Bearer ${OMNISTRATE_API_KEY_OR_TOKEN}" \ --header "Accept: application/json" ``` ### Example response ``` { "endpoint": "https://metrics.example.com/v1/customUsageMetering/", "expiresAt": "2026-09-23T00:00:00Z", "token": "eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9.example.signature" } ``` - `endpoint`: The complete URL for submitting custom metrics. Use it exactly as returned; do not construct or append the API path yourself. - `token`: A temporary bearer token for sending custom metrics for your organization. - `expiresAt`: The token expiration time. Refresh the endpoint credentials before this RFC 3339 timestamp. ### Load the response into shell variables ``` CREDENTIALS="$(curl --silent --show-error --fail \ "https://api.omnistrate.cloud/2022-09-01-00/custom-metrics/endpoint?expiresInSeconds=86400" \ --header "Authorization: Bearer ${OMNISTRATE_API_KEY_OR_TOKEN}" \ --header "Accept: application/json")" export METRICS_ENDPOINT="$(printf '%s' "${CREDENTIALS}" | jq -r '.endpoint')" export METRICS_TOKEN="$(printf '%s' "${CREDENTIALS}" | jq -r '.token')" export METRICS_TOKEN_EXPIRES_AT="$(printf '%s' "${CREDENTIALS}" | jq -r '.expiresAt')" ``` ## Send a Custom Usage Event Send an HTTP `POST` request to the exact URL returned in `endpoint`. Authenticate with the temporary metrics token, not the API key or token used to obtain the credentials. Note The request body is always an array. Even when sending one event, enclose the event object in square brackets. The body must not exceed 4 MiB and may contain at most 100 events. Replace the sample timestamps below with a timestamp from the current or previous UTC hour. Important Neither `subscription-id` nor `instance-id` is individually mandatory. Every event must contain at least one of them. If you provide only `instance-id`, Omnistrate resolves the subscription automatically. ### Use only a subscription ID ``` curl --silent --show-error --request POST \ "${METRICS_ENDPOINT}" \ --header "Authorization: Bearer ${METRICS_TOKEN}" \ --header "Content-Type: application/json" \ --data '[ { "idempotency-id": "usage-event-001", "subscription-id": "subscription-123", "timestamp": "2026-08-23T14:30:00Z", "metrics": { "requests": 1250, "storage-bytes": 4096 } } ]' ``` ### Use only an instance ID If you provide only an `instance-id`, the API resolves its subscription automatically. ``` [ { "idempotency-id": "usage-event-002", "instance-id": "instance-123", "timestamp": "2026-08-23T14:31:00Z", "metrics": { "requests": 42 } } ] ``` The Resource instance must already be associated with a subscription. Omnistrate determines that subscription and records the usage against it. Immediately after registering an instance, it may take a short time before the instance is available for usage reporting. If the event fails with `instance not found`, retry it with the same `idempotency-id`. ### Use both identifiers You may provide both `subscription-id` and `instance-id`. Both identifiers must refer to the same subscription. ``` [ { "idempotency-id": "usage-event-003", "subscription-id": "subscription-123", "instance-id": "instance-123", "timestamp": "2026-08-23T14:32:00Z", "metrics": { "requests": 100 } } ] ``` ### Required event fields - `idempotency-id`: A non-empty client-generated retry key, scoped to the authenticated organization and limited to 256 characters. - `timestamp`: An RFC 3339 timestamp within the previous or current UTC clock hour. The service stores it in UTC with second precision and uses it to determine the UTC billing period. - `metrics`: A non-empty object containing at most 100 metric names mapped to signed 64-bit integers. Metric names are limited to 256 characters; decimal values are rejected. - Target: Include `subscription-id`, `instance-id`, or both. Each identifier is limited to 256 characters. When only `instance-id` is supplied, Omnistrate determines the subscription. When both are supplied, they must refer to the same subscription. Warning Preserve integer precision. If your programming language cannot safely represent the full signed 64-bit integer range, use a JSON library or numeric type that preserves integer values exactly. ### Registration response A processed request returns `200 OK` with separate `successful` and `failed` arrays. Always inspect both arrays because a `200` response can contain rejected events. ``` { "successful": [ { "event-id": "0c4596cb-b85f-4e6d-bdb8-4cd16ed95b3e", "idempotency-id": "usage-event-001", "subscription-id": "subscription-123", "timestamp": "2026-08-23T14:30:00Z", "billing-period": "20260823", "metrics": { "requests": 1250, "storage-bytes": 4096 }, "created-at": "2026-08-23T14:31:00Z", "updated-at": "2026-08-23T14:31:00Z", "version": 1 } ], "failed": [] } ``` The service generates `event-id`, `billing-period`, `created-at`, `updated-at`, and `version`. Do not send or depend on client-provided values for these fields. ## Send Multiple Events in One Request Place multiple event objects in the same JSON array. You can mix subscription-based and instance-based events. ``` [ { "idempotency-id": "usage-event-004", "subscription-id": "subscription-123", "timestamp": "2026-08-23T14:32:00Z", "metrics": { "requests": 100 } }, { "idempotency-id": "usage-event-005", "instance-id": "instance-456", "timestamp": "2026-08-23T14:32:00Z", "metrics": { "storage-bytes": 4096 } } ] ``` ### Batch failure behavior - A syntactically malformed body, invalid JSON field type, empty batch, or request-limit violation rejects the complete request before storage with `400 Bad Request` or `413 Payload Too Large`. - An event that is valid JSON but fails field, timestamp, subscription, instance, metric-definition, or idempotency validation is added to `failed`. Other valid events in the batch continue processing. - Existing events are never updated. An `idempotency-id` that already exists with identical content is added to `successful`; different content is added to `failed` and leaves the stored event unchanged. - Target and idempotency throughput controls add affected events to `failed`. They do not change the HTTP status or add a `Retry-After` header. - If an infrastructure failure occurs after some events were stored, the request returns `500 Internal Server Error` without a partial result. Retry the complete batch with the same idempotency IDs; previously stored events return their original records and the remaining events are attempted again. ## Retries, Idempotency, and Throughput ### Safe retry behavior - Retry an event with the same `idempotency-id` and identical content. - After the one-second retry window, an identical retry returns the original event, including its `event-id` and timestamps. - Reusing an ID with different content adds the event to `failed` in the `200 OK` response. - Existing events are never updated. ### Throughput limits The API applies these limits independently for your organization: - Up to 10 HTTP requests per second. Exceeding this request-level limit returns `429 Too Many Requests` with `Retry-After: 1` before any event is processed. - Up to one accepted HTTP request per second containing the same target, with only one such request processed at a time. - Up to one accepted HTTP request per second containing the same `idempotency-id`, with only one such request processed at a time. For rate limiting, the target is the `instance-id` when present; otherwise it is the `subscription-id`. Requests for different targets can run concurrently within the organization-level limit. Target and idempotency limits return the affected events in `failed` as part of a `200 OK` response. Wait at least one second before retrying those events. ### Recommended retry logic 1. Keep the original event body and `idempotency-id`. 1. Inspect `failed` after every `200 OK`. Correct validation failures and wait at least one second before retrying target or idempotency throughput failures. 1. For `429`, wait for the number of seconds in `Retry-After`; no event in that request was processed. 1. For transient `500` responses, retry with exponential backoff and jitter. 1. For an expired metrics token, request fresh endpoint credentials and retry with the same event content and idempotency ID. ## Troubleshooting ### Getting the endpoint and token - `400 Bad Request`: `expiresInSeconds` is invalid or exceeds the configured maximum. - `401 Unauthorized`: The Omnistrate API key or token is missing or invalid. - `403 Forbidden`: The authenticated user does not have the Root, Admin, or Service Operator role required to retrieve custom metrics credentials. - `500 Internal Server Error`: Endpoint credentials could not be generated. Retry later or contact support. ### Sending custom metrics - `200 OK`: The body contains per-record results. Record validation, ownership, target, metric-definition, idempotency, and event-level throughput failures appear in `failed`. - `400 Bad Request`: The JSON is malformed, the array is empty or exceeds 100 events, an event exceeds 100 metrics, a field has an invalid JSON type, or an identifier or metric name exceeds 256 characters. - `401 Unauthorized`: The temporary metrics bearer token is missing, invalid, or expired. - `413 Payload Too Large`: The request body exceeds 4 MiB. - `429 Too Many Requests`: Your organization exceeded the limit of 10 HTTP requests per second before processing. - `500 Internal Server Error`: The request failed internally. Retry safely with the same idempotency IDs. - `504 Gateway Timeout`: Request processing exceeded the configured server timeout. Common per-record failures include: - `subscription not found`: Verify that the subscription exists in your organization and try again. - `instance not found`: Verify that the instance exists and is associated with a subscription. For a recently registered instance, wait briefly and retry with the same `idempotency-id`. - `instance does not belong to the specified subscription`: Both identifiers were supplied but do not match. - `timestamp must not be after instance deletion`: The event occurred after the resolved instance was deleted. ## Quick Reference - Credential endpoint: `GET https://api.omnistrate.cloud/2022-09-01-00/custom-metrics/endpoint` - Credential access: Root, Admin, or Service Operator. Authenticate with an Omnistrate API key or token. - Metrics endpoint: Use the complete URL returned in the credential response. - Authentication: `Bearer ` (send as the `Authorization` header value) - Content type: `application/json` - Request shape: A JSON array containing 1–100 events, with at most 100 metrics per event and a total body size no greater than 4 MiB. - Target: Include at least one of `subscription-id` or `instance-id`. Either identifier can be used alone, or both can be supplied together. - Event retention: Controlled by the service configuration; the default is seven days. For the complete API operation reference, see [Get custom metrics endpoint](https://api.omnistrate.cloud/docs/external/#tag/sp-organization-api/operation/sp-organization-api#GetCustomMetricsEndpoint). # Cloud Marketplace Integrations ## Cloud Marketplace Integrations Overview Cloud marketplace integrations enable software vendors to distribute and monetize their SaaS Product through major cloud provider marketplaces. These integrations provide a streamlined path to reach enterprise customers who prefer to procure software through their existing cloud provider relationships. ## What is a Cloud Marketplace Integration A cloud marketplace integration is a partnership mechanism that allows software vendors to list, sell, and distribute their applications through cloud provider marketplaces. These integrations offer several key benefits: - **Simplified Procurement**: Customers can purchase software using their existing cloud provider accounts and billing relationships - **Consolidated Billing**: Software costs appear on the customer's cloud provider bill, simplifying expense management - **Trust and Credibility**: Marketplace listings provide additional validation and trust for enterprise buyers - **Reduced Sales Friction**: Leverages existing customer relationships with cloud providers - **Global Reach**: Access to the cloud provider's worldwide customer base ## Major Cloud Marketplaces ### AWS Marketplace AWS Marketplace is Amazon's digital catalog that makes it easy for customers to find, buy, and immediately start using software and services that run on AWS. **Key Features:** - **Listing Types**: Software as a Service (SaaS), Amazon Machine Images (AMI), Containers, and Professional Services - **Pricing Models**: Free, Bring Your Own License (BYOL), Pay-as-you-go, Annual subscriptions, and Usage-based pricing - **Customer Base**: Millions of AWS customers across enterprise, government, and startup segments - **Payment Processing**: Handled through AWS billing with support for multiple currencies - **Private Offers**: Ability to create custom pricing and terms for specific customers **Benefits for Vendors:** - Access to AWS's extensive customer ecosystem - Streamlined procurement for enterprise customers - Comprehensive analytics and reporting tools ### Google Cloud Marketplace (GCP) Google Cloud Marketplace provides a platform for customers to discover, purchase, and deploy software solutions that run on GCP. **Key Features:** - **Solution Types**: SaaS applications, VM solutions, Kubernetes applications, and Developer tools - **Pricing Models**: Free trials, Pay-as-you-go, Subscription-based, and Bring Your Own License - **Customer Reach**: Access to Google Cloud's growing enterprise customer base - **Private Catalog**: Enterprise customers can create private catalogs with approved solutions **Benefits for Vendors:** - Leverage Google's enterprise relationships - Simplified customer onboarding and billing - Access to Google's partner ecosystem ### Microsoft Azure Marketplace Azure Marketplace is Microsoft's online store for applications and services certified to run on Azure, offering solutions from Microsoft and its partners. **Key Features:** - **Offer Types**: SaaS applications, Virtual Machines, Solution Templates, Managed Applications, and Consulting Services - **Pricing Models**: Free, Free trial, BYOL, Pay-as-you-go, and Monthly/Annual subscriptions - **Customer Base**: Extensive enterprise customer base with strong Microsoft relationships - **Co-sell Opportunities**: Access to Microsoft's sales team and partner programs - **Private Offers**: Custom pricing and terms for enterprise customers **Benefits for Vendors:** - Access to Microsoft's enterprise customer relationships - Co-selling opportunities with Microsoft sales teams - Comprehensive partner support programs ## Omnistrate Marketplace Support Omnistrate simplifies the process of creating and managing cloud marketplace listings across all major cloud providers. Our platform provides: ### Automated Listing Creation - **Multi-Cloud Support**: Create listings simultaneously across AWS, GCP, and Azure marketplaces - **Standardized Process**: Unified workflow for marketplace onboarding regardless of cloud provider - **Compliance Management**: Ensure listings meet each marketplace's technical and business requirements ### Integration Management - **Billing Integration**: Integration with cloud provider billing systems - **Metering and Usage Tracking**: Usage reporting compatible with marketplace billing systems ## Example Use Case: Integrating Billing with Marketplace Learn how to integrate your SaaS Product with Marketplace through our comprehensive step-by-step guide. This example covers: - Enabling tenant billing and configuring billing providers - Setting up usage metering and data export to cloud storage - Managing customer subscriptions and marketplace billing - Analyzing consumption data with Amazon Athena - Automating usage reporting to Cloud Marketplace using Clazar Demo Video: For a concrete example of exporting Omnistrate metered usage into a Clazar contract for cloud marketplace billing, see the [usage-export-clazar-recipe](https://github.com/omnistrate-community/usage-export-clazar-recipe). ## Evaluate your Marketplace integration needs For more information about cloud marketplace integrations and how Omnistrate can help you succeed in cloud marketplaces, please reach out to our support team at [support@omnistrate.com](mailto:support@omnistrate.com) Our team of marketplace experts is ready to help you navigate the complexities of cloud marketplace integrations and maximize your success across AWS, GCP, and Azure marketplaces. # Measure your SaaS Product usage ## Metering Overview Alternatively, a SaaS Provider may enable "Omnistrate Metering" to collect usage data and integrate with your billing system. The usage data will contain a line for each billing dimension and will include information about the cloud, region, and customer that used the service. The information for all your clouds and services that you have is collected in a single S3 or GCS bucket, allowing you to have a centralized view of the usage and process the billing information. Usage metering data from Omnistrate can be integrated with your existing billing system or used to manually generate invoices. The usage data generated by Omnistrate might require custom transformations as mandated by your billing provider. Omnistrate provides a data structure that allows to perform the required transformations in a simple way. To report application-specific usage directly to Omnistrate, see [Send Custom Usage Metrics](https://docs.omnistrate.com/fin-ops-guides/custom-usage-metering/index.md). ## Metering as the integration boundary for external billing For BYO billing providers and cloud marketplace adapters, the S3 or GCS export is the handoff point between Omnistrate and your downstream billing workflow. Omnistrate normalizes and exports the usage data. Your downstream system is responsible for: - Reading the exported files - Transforming the records into the schema expected by your billing provider or marketplace - Submitting usage or generating invoices outside Omnistrate If you need to correlate exported usage with an external customer or payer record, pass an `externalPayerId` during instance creation. Omnistrate includes that value in every exported usage record for the instance. Tip Configure Billing event alarms such as `S3MeteringExportFailed` or `GCSMeteringExportFailed` so export issues are detected quickly. See [Alarms](https://docs.omnistrate.com/operate-guides/alarms/index.md). To enable metering for your SaaS, follow these steps: 1. Grant privileges to Omnistrate to write to your S3/GCS bucket 1. Enable metering export on your Plan from UI or your compose spec 1. Using the UI, navigate to the *"Build Service > Plans -> Modify Plan"* and configure the bucket name in the `Metering` section. 1. Using the compose spec, provide the `metering` specifications under the `x-omnistrate-service-plan` section in your [Compose Specification](https://docs.omnistrate.com/build-guides/compose-spec/index.md). 1. Using the plan spec, provide the `metering` specifications in your [Plan Specification](https://docs.omnistrate.com/spec-guides/plan-spec/index.md). Here is an example of the metering configuration for your Plan: ``` metering: s3BucketARN: arn:aws:s3:::my_billing_bucket_name s3BucketRegion: us-west-2 gcsBucketName: my-billing-bucket-name ``` - `metering`: The metering configuration of the Plan. You can meter your infrastructure to capture the usage per customer, aggregate and store them at a defined location in your account. For a detailed step-by-step guide on configuring your bucket with the appropriate policy for storing metering data, please see [here](https://docs.omnistrate.com/fin-ops-guides/billing/index.md). - `s3BucketARN`: The ARN of the S3 bucket where the metering data will be stored. - `s3BucketRegion`: The region of the S3 bucket. e.g. `us-east-1`, `us-west-2`, etc. - `gcsBucketName`: The name of the GCS bucket where the metering data will be stored. ## Enabling metering integration for S3 (AWS) 1. Create a bucket in S3 To create a bucket in Amazon S3, you need to log in to the AWS Management Console, navigate to the S3 service, and click on the "Create bucket" button. In the dialog that appears, provide a unique name for your bucket, select the AWS region where the bucket will be hosted, and configure the options such as versioning, encryption, and access control settings. Once you have configured the desired options, click "Create" to finalize the process. The newly created bucket will now be available for uploading files and managing your data. You can also create a bucket using the AWS CLI or SDKs. Note If there is no location restriction we recommend using us-west-2 as bucket location. 1. Grant privileges To set a bucket policy in Amazon S3, navigate to the S3 console, select the desired bucket, and go to the "Permissions" tab. Under the "Bucket Policy" section, you can add a JSON policy to define access permissions for the bucket. This policy specifies who can access the bucket, what actions they can perform, and on which resources. After writing the policy, click "Save" to apply it. You can also use the AWS CLI or SDKs. ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::498789612402:root" }, "Action": [ "s3:PutObject", "s3:GetObject" ], "Resource": "arn:aws:s3:::[bucketName]/*" } ] } ``` Note 498789612402 is the service account used by Omnistrate to export the data. 1. Configure bucket name in your Product Offering You can add the configuration on the Plan through Omnistrate API/UI. Only the bucket ARN is required. ## Enabling metering integration for GCS (GCP) 1. Create a bucket in GCS To create a bucket in Google Cloud Storage, log into the Google Cloud Console, go to the "Storage" section, and click "Create bucket." Choose a unique name, select a storage class and region, and configure any additional settings. Click "Create" to finish. Alternatively, you can use Google Cloud CLI or the Google Cloud APIs to create the bucket. 1. Grant privileges To grant privileges to Omnistrate to write to the bucket you need to manage the permissions on the bucket and grant **Storage Object Admin** to omnistrate-billing@omnistrate-prod.iam.gserviceaccount.com Note omnistrate-billing@omnistrate-prod.iam.gserviceaccount.com is the Service Account used by Omnistrate to write to the bucket. You can restrict access only to the required bucket and provide **Storage Object Admin** privilege only for that bucket. 1. Configure bucket name in your Product Offering You can add the configuration on the Plan through Omnistrate API/UI. Only the bucket name is required. ## Enabling metering integration for Azure Blob Storage Exporting billing data to Azure is not yet available. Please reach out to [support@omnistrate.com](mailto:support@omnistrate.com) if you need it for your use case. ## Data structure Each exported file contains an array of JSON objects, each representing metering data for a specific instance. ``` [ { "timestamp": "2025-02-27T08:00:16Z", "organizationId": "org-xxx", "customerId": "user-xxx", "organizationName": "", "customerEmail": "xx@asdasd.com", "subscriptionId": "sub-xxx", "externalPayerId": "xxx", "serviceId": "s-xxx", "serviceName": "example-service", "serviceEnvironmentId": "se-xxx", "serviceEnvironmentType": "Dev", "productTierId": "pt-xxx", "productTierName": "Basic", "hostClusterId": "hc-79g18pwru", "cloudProvider": "aws", "region": "ca-central-1", "customNetworkId": "network-custom123", "instanceId": "instance-dxe51cv37", "podName": "rediscluster-replicas-0", "instanceType": "t4g.small", "hostName": "ip-172-0-33-151.ca-central-1.compute.internal", "dimension": "metricA", "value": 1, "pricePerUnit": 1.00 }, { "timestamp": "2025-02-27T08:00:16Z", "organizationId": "org-xxx", "subscriptionId": "sub-xxx", "customerId": "user-xxx", "customerEmail": "xx@asdasd.com", "externalPayerId": "xxx", "serviceId": "s-xxx", "serviceName": "example-service", "serviceEnvironmentId": "se-xxx", "serviceEnvironmentType": "Dev", "productTierId": "pt-xxx", "productTierName": "Basic", "hostClusterId": "hc-79g18pwru", "cloudProvider": "aws", "region": "ca-central-1", "instanceId": "instance-dxe51cv37", "podName": "rediscluster-master-0", "instanceType": "t4g.small", "hostName": "ip-172-0-33-151.ca-central-1.compute.internal", "dimension": "metricB", "value": 2, "pricePerUnit": 0.00 }, { "timestamp": "2025-02-27T08:00:16Z", "organizationId": "org-xxx", "subscriptionId": "sub-xxx", "customerId": "user-xxx", "customerEmail": "xx@asdasd.com", "externalPayerId": "xxx", "serviceId": "s-xxx", "serviceName": "example-service", "serviceEnvironmentId": "se-xxx", "serviceEnvironmentType": "Dev", "productTierId": "pt-xxx", "productTierName": "Basic", "hostClusterId": "hc-79g18pwru", "cloudProvider": "aws", "region": "ca-central-1", "instanceId": "instance-dxe51cv37", "meteringId": "pvc:data-rediscluster-master-0", "podName": "pvc:data-rediscluster-master-0", "storageClassName": "gp3", "volumeType": "AWS::EBS_GP3", "dimension": "storage_allocated_byte_hours", "value": 10737418240, "pricePerUnit": 0.02 } ] ``` ## Storage Path Format Files are stored per subscription in the following folder structure: ``` /omnistrate-metering/{service-name}/{environment}/{service-plan-id}/{year}/{month}/{day}/{hour}/{subscription-id}.json ``` Example: ``` /omnistrate-metering/example-service/dev/pt-123/2025/02/27/08/sub-xxx.json ``` There is a single file per subscription per hour, but the file will be updated before the end of the hour in case of suspension. ## Tracking Export Status Omnistrate maintains an **export status file** at: ``` /omnistrate-metering/last_success_export.json ``` This JSON file records the **last successfully processed timestamp** for each service and environment combination. The file is updated when: - A new export of metering data completes successfully, **or** - No new data is found for a given service plan during the export run. ### File Structure The JSON file is a map where: - **Key** → A unique identifier in the form ``` :: ``` \* **Value** → An object containing: - **`last_processed_to`** – All metering data up to this timestamp (UTC) has been processed and exported for the service plan. - **`last_updated_at`** – The timestamp (UTC) when the export status was last updated. ### Example ``` { "example-service:DEV:pt-xxx": { "last_processed_to": "2025-08-08T09:44:59Z", "last_updated_at": "2025-08-12T01:55:01Z" }, "example-service:PROD:pt-xxx": { "last_processed_to": "2025-08-08T09:44:59Z", "last_updated_at": "2025-08-12T01:55:01Z" } } ``` ## Understanding the JSON File Structure Each exported JSON file contains an array of records, with one entry per metered resource per hour per dimension (metric). Compute records are usually scoped to a Kubernetes pod. Persistent storage records are scoped to a persistent volume claim (PVC). Below is a breakdown of each field: - timestamp – The latest update time for the recorded usage, in ISO 8601 format (YYYY-MM-DDTHH:MM:SSZ). - organizationId – A unique identifier for the organization that owns the subscription. - customerId – Represents the customer associated with the subscription. - organizationName – The name of the organization. - customerEmail – The email address of the customer associated with the subscription. - subscriptionId – A unique identifier for the subscription that the service is being used under. - externalPayerId – A reference to the external billing entity responsible for payments. This value can be passed optionally during the instance creation. - serviceId – A unique identifier for the service generating usage data. - serviceName – The name of the service. - serviceEnvironmentId – A unique identifier for the environment (e.g., Dev, Prod) where the service instance is running. - serviceEnvironmentType – The environment type, such as "Dev", "Staging", or "Production". - productTierId – A unique identifier for the product tier (e.g., Free, Basic, Enterprise) associated with the subscription. - productTierName – The name of the product tier. - hostClusterId – The ID of the host cluster where the instance is deployed. - cloudProvider – The normalized cloud provider identifier for the host cluster where the instance is deployed, such as `aws`, `gcp`, or `azure`. - region – The cloud provider's native region identifier for the host cluster where the instance is deployed, such as `us-east-1` for AWS or `us-central1` for Google Cloud. - customNetworkId – The ID of the custom network attached to the host cluster, when the instance is deployed in a custom network. This field is omitted when no custom network applies. - instanceId – A unique identifier for the specific instance running the service. This is a key field for identifying unique records. - meteringId – A stable identifier for the metered resource when the record is not strictly pod-scoped. PVC storage records use an identifier such as `pvc:`. This field is omitted for older records and for records that do not need a separate metering identity. - podName – The Kubernetes pod name for pod-scoped compute records. PVC storage records use `pvc:`. - instanceType – The type of instance running the service, such as "t4g.small". This field may be empty for PVC storage records because PVCs are not tied to a specific compute node. - hostName – The hostname of the instance, often referring to a cloud VM or compute node. This field may be empty for PVC storage records because PVCs are not tied to a specific compute node. - storageClassName – The Kubernetes StorageClass name for PVC storage records. This field is omitted for non-storage records. - volumeType – The cloud provider volume type for PVC storage records when Omnistrate can resolve it. This field is omitted for non-storage records and may be empty when the provider volume type cannot be resolved. - dimension – The specific metric being recorded (e.g., CPU usage, memory consumption, replica usage). - value – The recorded metric value, representing the maximum usage observed during the hour. - pricePerUnit – The price per unit for the recorded dimension, allowing for cost calculation based on usage. ### GPU metering records GPU-enabled workloads emit GPU usage when their pods request GPU resources. The raw usage event field is `gpuMillicoreCount`, where a full allocatable GPU pool on the scheduled node is normalized to 1000 millicores. Exported metering reports this as the dimension `gpu_millicore_hours`; in pricing configuration, use `dimension: gpu` with `unit: millicores` and `timeUnit: hour`. When aggregating to instance-level usage and invoices, Omnistrate uses the maximum GPU millicore value per instance for the hour (not the sum across pods). ### Storage metering records Persistent storage usage is metered from Kubernetes PVCs. Exported storage records use the dimension `storage_allocated_byte_hours`; in pricing configuration, use `dimension: storage` with `unit: GiB` and `timeUnit: hour`. PVC storage records use `meteringId` and `podName` values such as `pvc:` so downstream billing systems can identify the storage resource independently from any pod that mounts it. Because PVCs are storage resources rather than compute resources, `instanceType` and `hostName` may be empty on storage records. When available, Omnistrate includes `storageClassName` and `volumeType` to describe the underlying storage class and cloud volume type. ### Processing the Exported Data Customers can process the exported JSON files using: - Data Warehouses: Load data into BigQuery, Snowflake, or Redshift for analysis. - ETL Pipelines: Process data using Apache Spark, AWS Glue, or Google Dataflow. - Billing Systems: Integrate with financial reporting tools for customer invoicing. ### Querying and Analyzing Exported Data Based on Folder Structure The exported data is stored in a structured folder hierarchy that allows for easy aggregation and querying. Understanding the storage path format helps in organizing and retrieving data efficiently. Files are stored per subscription in the following folder structure: ``` /omnistrate-metering/{service-name}/{environment}/{service-plan-id}/{year}/{month}/{day}/{hour}/{subscription-id}. json ``` 1. View Usage Per Subscription To get a full view of a subscription's usage over time, you can retrieve all files under: ``` /omnistrate-metering/{service-name}/{environment}/{service-plan-id}/{year}/{month}/{subscription-id}.json ``` Approach: Aggregate all JSON files for a subscription within the desired time range. Example: To get usage for a subscription in February 2025: - Scan files from /2025/02/ - Aggregate value fields for relevant dimensions. 1. Monthly Compute & Memory Usage Summary To compute total CPU and Memory usage across all subscriptions in a given month, process all files under: ``` /omnistrate-metering/{service-name}/{environment}/{service-plan-id}/{year}/{month}/ ``` Approach: - Read all files in /2025/02/ to cover the entire month. - Sum up value for dimensions like CPU and Memory. 1. View Usage Per Organization & Product Tier To analyze usage at the organization or product tier level, scan across all subscriptions and group by: ``` /omnistrate-metering/{service-name}/{environment}/{service-plan-id}/{year}/{month}/{subscription-id}.json ``` Approach: - Read all files under /2025/02/ to get full monthly data. - Aggregate usage based on organizationId and productTierName. By structuring queries around the folder hierarchy, customers can efficiently retrieve and process usage data for subscriptions, compute & memory consumption, among other metrics, and organizational breakdowns. ## External Billing ID For more complex scenarios, your billing system may require adding specific customer billing details to usage records. SaaS Providers can accommodate this by assigning an [external billing id](https://api.omnistrate.cloud/docs/external/#tag/resource-instance-api/operation/resource-instance-api#CreateResourceInstance) when creating resource instances for customers. To configure it in the UI, navigate to *"FinOps Center > Tenant Pricing -> Modify Tenant Pricing"* and set the `External Payer ID` field. When set, this value will be included in the exported metering data under the `externalPayerId` field. # FinOps with Omnistrate Omnistrate provides comprehensive financial operations (FinOps) capabilities to help you manage, monitor, and optimize the financial aspects of your SaaS business. From usage-based billing to cost optimization, our platform offers the tools you need to maximize profitability while maintaining operational efficiency. ## FinOps Capabilities ### Complete Billing Management **[End-to-End Billing](https://docs.omnistrate.com/fin-ops-guides/billing/index.md)** provides a comprehensive solution for SaaS providers who want to manage their entire customer billing workflow through Omnistrate. This approach eliminates the need for separate billing infrastructure and provides: - **Automated Usage Metering**: Collect and aggregate usage data across all nodes and customers automatically - **Flexible Pricing Configuration**: Set up pricing plans with CPU, memory, storage and replica dimensions, plus custom quotas - **Integrated Payment Processing**: Connect with Stripe or other payment providers for seamless payment collection - **Invoice Management**: Automatic invoice generation, notifications, and payment tracking - **Customer Portal**: Self-service portal where customers can view usage, manage payments, and track invoices This solution is ideal for SaaS providers who want to focus on their core product while leveraging Omnistrate's proven billing infrastructure. ### Bring Your Own Billing (BYOB) **[Usage Metering](https://docs.omnistrate.com/fin-ops-guides/metering/index.md)** enables SaaS providers with existing billing systems to collect detailed usage data from Omnistrate and integrate it with their current financial operations. This approach provides: - **Granular Usage Data**: Export detailed usage metrics per metered resource, per hour, with comprehensive metadata - **Flexible Export Options**: Send data to S3, GCS, or other storage systems for processing - **Structured Data Format**: JSON format with standardized fields for easy integration - **Real-time Metering**: Continuous usage tracking with hourly data exports - **Custom Billing Integration**: Support for external billing IDs and custom customer identifiers - **Multi-cloud Support**: Unified usage data across AWS, GCP, and Azure deployments This approach allows organizations to maintain their existing billing relationships and processes while leveraging Omnistrate's infrastructure metering capabilities. ### Marketplace Integration **[Cloud Marketplace Integrations](https://docs.omnistrate.com/fin-ops-guides/marketplace/index.md)** help SaaS providers expand their reach and simplify customer acquisition through major cloud provider marketplaces. This capability offers: - **Multi-Cloud Marketplace Support**: List your SaaS on AWS Marketplace, Google Cloud Marketplace, and Azure Marketplace - **Streamlined Procurement**: Enable customers to purchase through their existing cloud provider relationships - **Consolidated Billing**: Customers receive unified bills through their cloud provider accounts - **Enterprise Trust**: Leverage cloud provider validation and security certifications - **Global Distribution**: Access worldwide customer bases through established marketplace channels - **Private Offers**: Create custom pricing and terms for enterprise customers Marketplace integration reduces sales friction and provides access to enterprise customers who prefer to procure software through their cloud provider relationships. ### Cost Control & Insights **[Cost Control & Insights](https://docs.omnistrate.com/fin-ops-guides/cost-insights/index.md)** provides comprehensive visibility and optimization tools for managing infrastructure costs across your SaaS deployments. This feature delivers: - **Real-time Cost Monitoring**: Track infrastructure costs across environments, deployments, and resources - **Granular Cost Attribution**: Use infrastructure tags to categorize spending by tenants - **Optimization Recommendations**: Identify opportunities for cost reduction through spot instances, reserved capacity, and auto-scaling - **Automated Cost Controls**: Implement auto-stop and auto-scaling policies to prevent cost overruns These insights enable data-driven decisions about infrastructure spending while maintaining performance and reliability standards. # Use Cases # Agent as a Service (AaaS) Agent as a Service (AaaS) is a cloud computing model that allows you to build, deploy, and manage AI agents without the complexity of managing the underlying infrastructure. With Omnistrate, you can create production ready Agents that are scalable, multi-tenant and can be delivered to customers on any cloud and even to your customer account (BYOC). ## Why Build AI Agents with Omnistrate? Omnistrate provides a comprehensive platform to operationalize your AI investments and launch your own Agent as a Service. Here’s how you can benefit: - **Accelerate Time-to-Market**: Go from a single-tenant AI application to a full-fledged, multi-tenant SaaS product in weeks, not months. - **Reduce Operational Complexity**: Offload the complexities of infrastructure management, multi-tenancy, and cross-cloud deployment to a single, unified platform. - **Scale with Confidence**: Omnistrate’s architecture is designed for high availability and scalability, ensuring your AI agents can handle growing demand. - **Retain Full Control**: Maintain control over your data, IP, and customer relationships while leveraging the power of a managed platform. ## How to Create AI Agents with Omnistrate Building an AI agent with Omnistrate involves a few key steps: 1. **Package Your Application**: Containerize your AI model, business logic, and any other dependencies into a container image. 1. **Define Your Plan Blueprint**: Create a Plan that specifies the resources, configurations, and tenancy model for your AI agent. This defines how your agent will be deployed and managed. 1. **Onboard to Omnistrate**: Use the Omnistrate Customer Portal to onboard your Plan and configure your SaaS Product details. 1. **Deploy to Any Cloud**: With a single click, deploy your AI agent to any of the supported cloud providers (AWS, Azure, GCP) or even on-premises environments. Omnistrate handles the provisioning and configuration of all necessary resources. 1. **Manage and Monitor**: Use the Omnistrate platform to manage tenant lifecycle, monitor agent performance, and gain insights into usage and costs. ## Key Features for AaaS Omnistrate offers several features that are critical for building and managing a successful Agent as a Service that are specific to agents: - **Multi-tenant by Design**: Each customer gets isolated agent environments with their own data silos and configurations. Imagine the traditional PaaS-like deployments that take multiple micro-services and isolate them by tenant. The same applies to the agentic world especially when these apps involve more than just the agentic app. - **Orchestration of Complex Architectures:** Agentic apps are increasingly becoming complex involving multiple components such as Vector stores, Memory stores, Custom LLMs hosted on specialized hardware, MCP servers, Tool uses, Observability and Tracing, Outcome-based Billing, etc. This leads to the agentic app developer stitching together a bunch of integrations just to get started. This is just to get something working for one user. - **Customer-Controlled Environments**: Agents deploy into customer's own cloud accounts (BYOA) for maximum security or save cost and deploy in isolated environments on the service provider’s accounts - **Self-service Deployments:** End customers can customize their deployments to tailor to their needs through traditional parametrization based on the Service Provider’s configuration and chosen flexibility - **Simplified Agent Publishing** - **Framework-Agnostic**: Works with OpenAI Agents SDK, LangGraph, LangChain, or any custom framework - **Declarative Configuration**: Simple YAML files define the entire agent service stack - **Automated Infrastructure**: Vector stores, monitoring, and integrations are provisioned automatically ## Get Started Ready to build your own Agent as a Service? [Contact us](mailto:support@omnistrate.com) to learn more about how Omnistrate can help you turn your AI innovations into market-ready products. # Air-Gapped Deployments Air-gapped deployments are essential for customers who require their data and applications to remain within their own data centers, often due to strict security, data sovereignty, or regulatory compliance requirements. This model is common in sectors like finance, healthcare, and government, where control over the physical infrastructure is paramount. With Omnistrate, software providers can package their products as downloadable, self-contained installers. These installers are designed for easy deployment in a customer's environment, including air-gapped settings with no internet connectivity. ## Why Air-Gapped? Customers may require air-gapped solutions for several reasons: - **Data Sovereignty and Residency**: Data must be stored in a specific geographic location or country. - **Security and Compliance**: Strict internal security policies or regulatory mandates (like HIPAA, GDPR, or PCI-DSS) require full control over the infrastructure. This also includes environments that need to be compliant with standards like FedRAMP, which require isolated and highly controlled cloud environments. - **Isolated Cloud Accounts**: Customers may require the service to run in their own dedicated and isolated cloud accounts, separate from other tenants and the provider's own infrastructure. - **Air-Gapped Environments**: For maximum security, some environments are completely disconnected from the public internet. - **Low Latency**: Applications that require near-instantaneous response times may need to be located physically close to the end-users or other systems. - **Control over Infrastructure**: Full control over hardware, software, and network configurations. ## How It Works with Omnistrate Omnistrate enables software providers to support air-gapped customers by generating an installer package. This package contains everything needed to deploy and run the product in an isolated environment. The process is as follows: 1. **Package**: The software provider uses Omnistrate to package the application, its dependencies including container image binaries and packages, and configuration into a single installer. 1. **Distribute**: The installer is made available for download to the customer. 1. **Install**: The customer's operator downloads the installer and runs it in their air-gapped environment. The installation can be fully scripted and automated. 1. **Operate**: Once installed, the product runs self-sufficiently within the customer's infrastructure. Updates and patches are delivered as new versions of the installer. This approach allows software providers to reach a broader market, including enterprises with stringent IT requirements, without the need to build a separate, dedicated air-gapped version of their product from scratch. ## Common Challenges Common challenges in this model include: - Packaging all the components together as one unit - Generate 1-click install script for your customers with troubleshooting assistant - Deployment templatization for your customers to customize deployments - Managed upgrades with version management - Tenant Database to store deployment configuration, version, licensing, pricing, operational context - Enforce licensing to prevent unauthorized use - Collect troubleshooting data locally ## Benefits of Using Omnistrate for Air-Gapped Deployments Omnistrate simplifies the complexity of managing air-gapped software. By standardizing the packaging and deployment process, it reduces the engineering overhead for software providers. This means you can offer an air-gapped solution without diverting focus from your core product. Omnistrate also provides a consistent experience for your customers, whether they are deploying in the cloud or on their own hardware. ## Inventory and Configuration Management Even in disconnected environments, Omnistrate helps you maintain a clear view of your deployed instances. This allows you to track which versions are deployed where, manage configurations centrally, and ensure that all instances are up-to-date and compliant with your standards. ## Support and Diagnostics Supporting air-gapped installations can be challenging. Omnistrate makes it easier by providing tooling for remote support and diagnostics. Customers can grant temporary, secure access to their environment, allowing your support teams to troubleshoot issues without compromising the security of the air-gapped installation. Diagnostic bundles can be securely uploaded to the provider, giving you the insights needed to resolve problems quickly. ## Get Started Ready to build an air-gapped installer? Follow the step-by-step [Air-Gapped Installer Plan](https://docs.omnistrate.com/build-guides/air-gapped-helm-charts/index.md) guide to create your first installer specification. # BYOC On-Premise BYOC On-Premise lets a customer bring an existing Kubernetes cluster and use it as the deployment target for your Product. Omnistrate treats this target as the `byoc-onprem` provider with the `on-prem` deployment region. The customer owns the Kubernetes cluster, nodes, storage, network routing, and endpoint exposure. Your generated control plane operates deployments in that cluster through your product dataplane agent. Note `byoc-onprem` is the value you pass to the `--cloud-provider` flag in `omnistrate-ctl` (for example `omnistrate-ctl instance create --cloud-provider=byoc-onprem`). BYOC On-Premise is configured as a BYOC deployment Plan. Set up the provider side the same way you would for other BYOC Plans, then pass `--cloud-provider=byoc-onprem` when creating customer onboarding instances and product deployments. For the shared BYOC flow, see [BYOC cloud accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md). For the Plan schema, see [deployment schema](https://docs.omnistrate.com/spec-guides/plan-spec/#deployment-schema). This mode is different from an air-gapped installer. In BYOC On-Premise, the customer stays connected to your generated control plane. For air-gapped or fully disconnected environments, use the [Air-Gapped installer](https://docs.omnistrate.com/usecases/air-gapped/index.md) model instead. ## Architecture The product dataplane agent runs in the customer cluster and opens outbound mTLS/gRPC agent connections to your control plane. The same agent connection pattern is used for the Manager Service, Infra Metadata Manager, and Monitoring Service. The customer does not need to expose the Kubernetes API server publicly. The cluster needs outbound access to your control plane endpoints, plus any image registries and Helm chart registries required by your Plan and deployment-cell amenities. Application traffic is separate from control-plane traffic. The customer decides whether product endpoints are private-only, reachable through a corporate network, or exposed through a public IP or load balancer. ## Before You Start Confirm the customer cluster has the baseline requirements for your product: | Requirement | Details | | ------------------ | ---------------------------------------------------------------------------------------------------- | | Kubernetes cluster | A working Kubernetes cluster with a supported Kubernetes version | | `kubectl` access | The operator running the installation kit must target the intended cluster context | | Pod networking | CNI and pod-to-pod networking must work | | DNS and egress | Pods must resolve and reach your control plane endpoints and required registries | | Storage | Required StorageClasses must exist if the Plan uses persistent volumes | | Endpoint path | Ingress, load balancer, firewall, and DNS routing must match the endpoint exposure your Plan defines | If your product needs cluster-level components, configure them as deployment-cell amenities, or have the customer install equivalent prerequisites before deploying product instances. Examples include ingress controllers, cert-manager, monitoring, endpoint checks, logging, and custom Helm packages. The BYOC On-Premise install kit installs the dataplane agent; amenities and product workloads are handled after the agent connects. See [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md). ## End-to-End Walkthrough This walkthrough uses PostgreSQL as the example product. The same sequence applies to other products that support `byoc-onprem`. ### 1. Provider prepares the Plan Build and release a Plan that can deploy to customer-owned clusters. The product instance later selects BYOC On-Premise with `--cloud-provider=byoc-onprem`. ``` name: BYOC OnPrem Simple PG Test deployment: byoaDeployment: AwsAccountId: '' AwsBootstrapRoleAccountArn: 'arn:aws:iam:::role/omnistrate-bootstrap-role' services: - name: postgresChart network: ports: - 5432 endpointConfiguration: postgresEndpoint: host: "$sys.network.internalClusterEndpoint" ports: - 5432 primary: true networkingType: INTERNAL ``` BYOC On-Premise does not provision nodes; the customer-managed Kubernetes cluster is the deployment target. The Plan should define the product endpoints, parameters, persistent volumes, metrics, dashboards, and upgrade paths that apply to the customer deployment. This walkthrough uses an internal endpoint. ### 2. Create the BYOC On-Premise onboarding instance The customer onboarding instance represents one customer-owned Kubernetes cluster. Passing `--cluster-name` (instead of a cloud-provider account flag such as `--aws-account-id`) selects the BYOC On-Premise path. Create a working directory first. The onboarding command automatically downloads the generated install kit into the current directory. ``` mkdir -p dp-install-kit cd dp-install-kit omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --cluster-name= \ --cluster-description="Customer production Kubernetes cluster" ``` If you create the onboarding instance on behalf of a customer, add `--customer-email=`. Save the returned customer onboarding instance ID, for example `instance-abc123`. The command does not wait for the onboarding instance to become `READY`. After the dataplane agent connects, you pass the same onboarding instance ID as `--customer-account-id` in the deployment command. ### 3. Install the dataplane agent on the target cluster Extract the kit and confirm `kubectl` targets the intended customer cluster: ``` # If the kit was downloaded elsewhere, copy or move it into this directory first. tar xf byoc-onprem-install-kit-.tar kubectl config current-context ``` Run the installer against the current Kubernetes context: ``` ./install.sh --non-interactive ``` The installer creates the dataplane-agent resources in the cluster. Amenities and product workloads are installed later by your control plane. #### Test your deployment If you do not have a customer cluster yet and want to validate the flow end to end, the same install kit can spin up a local or single-node test cluster. For a fully local smoke test, create a k3d cluster: ``` ./install.sh \ --create-k3d-cluster \ --non-interactive ``` If your test host has a public IP and you want to validate a public endpoint Plan, create a single-node K3s cluster on that host and advertise the public IP: ``` PUBLIC_IP= ./install.sh \ --create-k3s-cluster \ --k3s-node-external-ip "${PUBLIC_IP}" \ --non-interactive ``` For a public endpoint Plan, use the external cluster endpoint in the Plan instead of the internal one: ``` services: - name: postgresChart network: ports: - 5432 endpointConfiguration: postgresEndpoint: host: "$sys.network.externalClusterEndpoint" ports: - 5432 primary: true networkingType: PUBLIC ``` In production, customers install into their existing Kubernetes cluster and skip the `--create-k3d-cluster` / `--create-k3s-cluster` flags. ### 4. Verify the account configuration After the agent connects to your control plane, the BYOC On-Premise account configuration moves to `READY`. ``` omnistrate-ctl account customer describe ``` Continue when `account_status` is `READY`. ### 5. Deploy a product instance Create the product instance and pass the BYOC On-Premise onboarding instance ID as `--customer-account-id`. For the PostgreSQL example, create a parameter file: ``` { "pgPassword": "change-me", "pgDatabase": "appdb" } ``` ``` omnistrate-ctl instance create \ --service= \ --environment= \ --plan= \ --version=latest \ --resource=postgresChart \ --cloud-provider=byoc-onprem \ --region=on-prem \ --customer-account-id= \ --param-file=./params.json \ --wait ``` Verify the instance from your control plane: ``` omnistrate-ctl instance describe omnistrate-ctl instance list-endpoints ``` Verify the workload in the customer cluster with the same Kubernetes context used to run `install.sh`. Restrict the query to the instance namespace: ``` kubectl get all -n helm list -n ``` ### 6. Operate the instance After the instance is running, lifecycle operations continue through your control plane: ``` omnistrate-ctl instance dashboard omnistrate-ctl instance version-upgrade --target-version= omnistrate-ctl instance delete ``` The customer can still inspect and operate their Kubernetes cluster with their own tools. Re-download the install kit for reinstall or recovery The normal onboarding flow does not need a separate download command because `omnistrate-ctl account customer create` automatically downloads the generated kit. To download the kit again for an existing customer onboarding instance, run: ``` omnistrate-ctl account customer install-kit ``` Then extract the fresh kit and rerun `./install.sh` against the target Kubernetes context. ## Native Logs BYOC On-Premise supports native log collection for Helm resources. When `features.INTERNAL.logs.provider: native` is set in a Helm-based Plan spec, Omnistrate deploys a per-instance OpenTelemetry collector that ships workload logs to CloudWatch Logs in the SaaS Provider's AWS account. ``` features: INTERNAL: logs: provider: native ``` No additional cloud logging credentials are required on the customer cluster beyond the standard dataplane agent installation. Omnistrate supplies the OpenTelemetry collector with scoped AWS credentials that write only to the instance's CloudWatch log groups and stream prefix in the SaaS Provider's AWS account. Keep `ExpandBoundaryForAWSServiceIntegrations=true` in that AWS account's generated CloudFormation stack so Omnistrate can provide these credentials. The customer cluster must have an HTTPS network path to the regional AWS CloudWatch Logs endpoint. It can use public egress or private, routed connectivity to an AWS CloudWatch Logs interface VPC endpoint. A private path requires appropriate routing and DNS, such as connectivity through a VPN or AWS Direct Connect. Native logs are not available in [air-gapped](https://docs.omnistrate.com/usecases/air-gapped/index.md) clusters that cannot reach CloudWatch Logs. For configuration details, see [Send Helm Resource Logs to CloudWatch Logs](https://docs.omnistrate.com/getting-started/build-from-helm/#send-helm-resource-logs-to-cloudwatch-logs). ## Limitations - The customer cluster must have outbound connectivity to your control plane. - BYOC On-Premise is not for fully air-gapped environments; use the [Air-Gapped installer](https://docs.omnistrate.com/usecases/air-gapped/index.md) for that model. - Each customer onboarding instance maps to one Kubernetes cluster. - Endpoint reachability depends on the customer's DNS, firewall, load balancer, and routing configuration. ## Related Guides - [BYOC cloud accounts](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/index.md) - [Deployment Cell Amenities](https://docs.omnistrate.com/operate-guides/deployment-cell-amenities/index.md) - [BYOC PrivateLink](https://docs.omnistrate.com/usecases/byoc-privatelink/index.md) - [Air-Gapped installer](https://docs.omnistrate.com/usecases/air-gapped/index.md) # BYOC PrivateLink Some customers cannot allow any public network exposure on the dataplane that runs your application — typical for regulated industries, strict data-residency setups, and customers with zero-public-egress policies. **BYOC PrivateLink** is a variant of [BYOC](https://docs.omnistrate.com/usecases/byoc/index.md) where every byte of control-plane traffic between the customer's dataplane and your Omnistrate control plane flows over **AWS PrivateLink**. The dataplane EKS cluster has no public endpoint and no internet-facing load balancers for control traffic. Note BYOC PrivateLink is currently supported on **AWS**. The customer's VPC and your provisioner VPC may live in different AWS regions. ## How to Enable BYOC PrivateLink BYOC PrivateLink is selected per **customer account**, at account onboarding time. Once an account is onboarded with PrivateLink enabled, every instance deployed into that account uses the PrivateLink dataplane topology — there is no per-instance toggle and no compose-spec change required. ### Customer Portal When your customer onboards their AWS account into a BYOC Plan from the customer portal, they enable the **PrivateLink connectivity** option on the AWS account form. The portal then issues the bootstrap CloudFormation link and tracks the account as PrivateLink-enabled. ### `omnistrate-ctl` ``` omnistrate-ctl account customer create \ --service= \ --environment= \ --plan= \ --customer-email= \ --aws-account-id= \ --private-link ``` The `--private-link` flag enables AWS PrivateLink connectivity for every instance deployed into this account. See the [`account customer create`](https://ctl.omnistrate.cloud/omnistrate-ctl_account_customer/) reference for all flags. ## BYOC PrivateLink Architecture The dataplane agent in the customer's account communicates with the Omnistrate-managed Control Plane through a customer-owned VPC Interface Endpoint that targets the PrivateLink service Omnistrate provides. No public internet is used for control traffic. When a PrivateLink account onboards for the first time, Omnistrate automatically configures the VPCE service and adds the customer's account to the service's `AllowedPrincipals` list. ## CloudFormation Controls BYOC PrivateLink accounts support the `K8sDebugAccessEnabled` CloudFormation control for Omnistrate support/debug access to the customer Kubernetes API. The same AWS account-config stack also includes `AgentInfrastructureMutationEnabled` for dataplane agent AWS infrastructure mutation permissions. For parameters, update steps, and verification, see [AWS CloudFormation Account Controls](https://docs.omnistrate.com/operate-guides/aws-cloudformation-account-controls/index.md). ## Customer VPC Topology Your customer can pick one of two VPC topologies for the dataplane: - **Omnistrate-managed VPC** — Omnistrate provisions the VPC, subnets, NAT gateway, security group, and the management VPCE on first deployment. Requires the **Allow new cloud-native network creation** option (or `--allow-create-new-cloud-native-network`) when the account is onboarded. See [Omnistrate-managed VPC](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#omnistrate-managed-vpc). - **Customer-owned (imported) VPC** — The customer provides an existing VPC and subnets, then creates the management VPCE. See [Imported VPC requirements for BYOC PrivateLink](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#imported-vpc-requirements-for-byoc-privatelink) for the exact VPC, subnet, and VPCE security group tags, security group ports, and cross-region constraints. # BYOC (Bring Your Own Cloud) There are many applications that needs to be deployed in customers account due to security and cost reasons. Your customers may prefer to not move the data in your account and want you to deploy your app(s) in their account. From your perspective, you will have to manage hundreds or thousands of these accounts. This is how the setup may look like: The challenge is that deploying in your customers account requires manual coordination, sharing of credentials and a lot of operational pain. We have automated the entire process and made it simple to operate in a secure way. Note There are several variants of BYOC mode in the industry and they are all somewhat related. - Bring Your Own VPC - in this mode, your customer brings a specific VPC for you to deploy and manage your application. See [here](#bring-your-own-vpc-byo-vpc) for more details - Bring Your Own Cloud (BYOC) - in this mode, your customers bring their account so that you can deploy and manage your application. See [here](#how-to-enable-byoc) for more details We support different variants of BYOC for you to NOT worry about the complexity of the underlying infrastructure ## How to enable BYOC ### Compose Spec Configuration If you are using compose-based specification, you can add the following to your compose to configure your provider account: ``` x-omnistrate-service-plan: deployment: byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role ``` Note Please don't forget to replace with your own AWS Account ID ### Videos - To configure your customer's account using Cloud Formation, you need to follow this [video guide](https://youtu.be/c3HNnM8UJBE) ## BYOC architecture We build a trust relationship between your account and your customers account to allow you to automate the setup. Once setup, system uses the industry standard secure techniques to reverse the connection to prevent any inbound connections to your customers' account (except while configuring their account during setup), encrypted channel through TLS and oauth tokens to secure the connectivity between your customers account and your account. If your customers wants to also disable any outbound data, they can also achieve that by updating the IAM permission set. Please reach out to [support@omnistrate.com](mailto:support@omnistrate.com) for details on how to achieve this. ## BYOC in action Omnistrate makes it easy to manage resource instances across the fleet ### View for your customers ### Internal view for your teams ### Demo video Here is a demo video on PostgreSQL BYOC DBaaS: ## Bring Your Own VPC (BYO-VPC) If you are running in BYOC mode, your customers can bring their own VPC and Omnistrate will deploy your Dataplane in their VPC. ### Prerequisites Your customer must create the VPC and subnets before deploying. The requirements differ depending on whether PrivateLink is enabled. #### Standard (PrivateLink disabled) | # | Requirement | Details | | --- | -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | **DNS settings** | Enable **DNS hostnames** and **DNS resolution** on the VPC. [AWS docs](https://docs.aws.amazon.com/vpc/latest/userguide/vpc-dns.html#vpc-dns-updating) | | 2 | **NAT Gateway** | A **public NAT gateway** is required for pulling container images. All **private subnet route tables** must have a route to this NAT gateway. [AWS docs](https://docs.aws.amazon.com/vpc/latest/userguide/WorkWithRouteTables.html#AddRemoveRoutes) | | 3 | **Public subnet auto-assign IP** | Public subnets must have **auto-assign public IPv4 address** enabled. [AWS docs](https://docs.aws.amazon.com/vpc/latest/userguide/modify-subnets.html#subnet-public-ip) | | 4 | **Subnet tags** | Private subnets: tag `kubernetes.io/role/internal-elb` = `1`. Public subnets: tag `kubernetes.io/role/elb` = `1`. [AWS docs](https://docs.aws.amazon.com/eks/latest/userguide/network-reqs.html#vpc-subnet-tagging) | #### PrivateLink enabled | # | Requirement | Details | | --- | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | **VPC & subnet tags** | Tag the VPC and workload subnets with `omnistrate.com/managed-by` = `omnistrate`. This tag gates IAM permissions and selects subnets for EKS nodegroup placement — tag **only** the subnets you want workloads in. | | 2 | **DNS settings** | Enable **DNS hostnames** (`enableDnsHostnames`) and **DNS resolution** (`enableDnsSupport`) on the VPC. | | 3 | **Egress** | Workload subnets need outbound internet via **NAT Gateway**, **Transit Gateway**, or **VPN**. Required for Helm binary downloads and container image pulls during dataplane bootstrap. | | 4 | **Management VPC Endpoint** | Create a single **Interface VPC Endpoint** targeting the PrivateLink service name Omnistrate provides after account onboarding, with: | | | | • Tag `Name` = `omnistrate-byoc-private-vpce-` | | | | • Security group allowing inbound TCP **8443–8506** from the VPC CIDR | | | | • Tag every security group attached to the VPCE with `omnistrate.com/managed-by` = `omnistrate`. Omnistrate uses this tag as the IAM permission gate when adding or removing Kubernetes debug access ingress rules. | | 5 | **Cross-region** *(if applicable)* | If the customer's VPC and the PrivateLink service are in different AWS regions, pass `--service-region ` when creating the VPC Endpoint (or `service_region` in Terraform). Do **not** enable private DNS — AWS does not support it for cross-region interface endpoints. | | 6 | **Subnet tags** *(optional)* | Private subnets: tag `kubernetes.io/role/internal-elb` = `1`. Public subnets: tag `kubernetes.io/role/elb` = `1`. Not mandatory for PrivateLink deployments, but recommended if you plan to use internal load balancers. | Note Omnistrate provides the VPCE service name and the provisioner host-cluster ID after the account is onboarded — both are needed to create the management VPC Endpoint. For full details on the PrivateLink VPC topology options, see [Imported VPC requirements for BYOC PrivateLink](https://docs.omnistrate.com/operate-guides/byoc-cloud-accounts/#imported-vpc-requirements-for-byoc-privatelink). ### How to get started When creating an instance, your customer can specify VPC id as value of `cloud_provider_native_network_id` input parameter, and rest of things will be same as regular BYOC experience. # Case Studies Omnistrate helps software teams remove the blockers that slow down enterprise software distribution: on-prem automation, multi-cloud launch, regulated customer deployments, internal delivery platforms, and day-2 operations. These customer stories show how teams use Omnistrate to ship deployable, operable SaaS Products while keeping engineering focused on the core product. ## Impact at a Glance | Company | What they used Omnistrate for | Impact | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | [DataRobot](https://blog.omnistrate.com/posts/datarobot-automates-on-prem-and-launches-sts?category=Customer+Stories) | Automated on-prem deployment and launched global multi-cloud STS | Saved 3 months of coding per customer and reduced on-prem deployment to a single step | | [Plotline](https://blog.omnistrate.com/posts/how-plotline-streamlines-internal-deployments-to-support-massive-scale-and-free-engineers-to-innovate?category=Customer+Stories) | Launched an internal delivery platform | Achieved 30%-50% TCO savings over alternatives and supported 1.5 billion daily API calls | | [FalkorDB](https://blog.omnistrate.com/posts/aws-and-omnistrate-partnered-to-bring-falkordb-managed-cloud-service?category=Customer+Stories) | Built a distribution control plane | Launched a secure, scalable SaaS in weeks and enabled thousands of global deployments with a single engineer | | [Experio](https://blog.omnistrate.com/posts/experio-solidified-its-ai-first-advantage-with-omnistrate?category=Customer+Stories) | Launched enterprise-ready multi-cloud SaaS | Saved 6+ months of engineering time and kept 100% focus on core AI innovation | | [TileDB](https://blog.omnistrate.com/posts/tiledb-accelerates-critical-enterprise-deals-by-automating-deployment-across-any-cloud-with-omnistrate?category=Customer+Stories) | Automated deployment for regulated industries | Launched an enterprise-ready Microsoft Azure offering in weeks with 100% isolation through dedicated, compliant per-tenant deployments | ## Customer Stories ### DataRobot: Automated on-prem and global STS DataRobot used Omnistrate to automate on-prem deployment and launch a global multi-cloud STS (single-tenant SaaS) offering that integrated with their existing stack. **Impact:** - Saved 3 months of coding per customer. - Reduced on-prem deployment to a single step. - Expanded deployment automation across on-prem and multi-cloud environments. [Read the DataRobot customer story](https://blog.omnistrate.com/posts/datarobot-automates-on-prem-and-launches-sts?category=Customer+Stories) ### Plotline: Internal delivery platform at massive scale Plotline used Omnistrate to build an internal delivery platform that replaced fragile manual processes with repeatable deployment automation. **Impact:** - Achieved 30%-50% TCO savings over alternatives. - Supported 1.5 billion daily API calls. - Freed engineering teams from dedicated delivery operations. [Read the Plotline customer story](https://blog.omnistrate.com/posts/how-plotline-streamlines-internal-deployments-to-support-massive-scale-and-free-engineers-to-innovate?category=Customer+Stories) ### FalkorDB: Multi-cloud distribution control plane FalkorDB used Omnistrate to build a distribution control plane for its GenAI database platform and launch a secure, scalable managed SaaS. **Impact:** - Launched a secure, scalable SaaS in weeks. - Enabled thousands of global deployments with a single engineer. - Accelerated multi-cloud availability without building the control plane from scratch. [Read the FalkorDB customer story](https://blog.omnistrate.com/posts/aws-and-omnistrate-partnered-to-bring-falkordb-managed-cloud-service?category=Customer+Stories) ### Experio: Enterprise-ready multi-cloud SaaS Experio used Omnistrate to move from pilot to enterprise-ready multi-cloud SaaS while preserving focus on its AI-first product experience. **Impact:** - Saved 6+ months of engineering time. - Kept 100% of engineering focus on core AI innovation. - Accelerated enterprise readiness for multi-cloud SaaS delivery. [Read the Experio customer story](https://blog.omnistrate.com/posts/experio-solidified-its-ai-first-advantage-with-omnistrate?category=Customer+Stories) ### TileDB: Regulated enterprise deployments TileDB used Omnistrate to automate deployment for regulated industries and launch an enterprise-ready offering on Microsoft Azure. **Impact:** - Launched an enterprise-ready Microsoft Azure offering in weeks. - Delivered 100% isolation through dedicated, compliant per-tenant deployments. - Streamlined enterprise customer onboarding through deployment automation. [Read the TileDB customer story](https://blog.omnistrate.com/posts/tiledb-accelerates-critical-enterprise-deals-by-automating-deployment-across-any-cloud-with-omnistrate?category=Customer+Stories) ## References - [Omnistrate Customer Stories](https://blog.omnistrate.com/?category=Customer+Stories) - [DataRobot Automates On-Prem and Launches STS](https://blog.omnistrate.com/posts/datarobot-automates-on-prem-and-launches-sts?category=Customer+Stories) - [How Plotline Streamlines Internal Deployments to Support Massive Scale and Free Engineers to Innovate](https://blog.omnistrate.com/posts/how-plotline-streamlines-internal-deployments-to-support-massive-scale-and-free-engineers-to-innovate?category=Customer+Stories) - [How FalkorDB Used Omnistrate for a Rapid Multi-Cloud Launch to Enter the GenAI Market](https://blog.omnistrate.com/posts/aws-and-omnistrate-partnered-to-bring-falkordb-managed-cloud-service?category=Customer+Stories) - [Experio Solidifies 6-Month AI Advantage with Multi-Cloud SaaS Launch](https://blog.omnistrate.com/posts/experio-solidified-its-ai-first-advantage-with-omnistrate?category=Customer+Stories) - [TileDB Accelerates Critical Enterprise Deals by Automating Deployment Across Any Cloud with Omnistrate](https://blog.omnistrate.com/posts/tiledb-accelerates-critical-enterprise-deals-by-automating-deployment-across-any-cloud-with-omnistrate?category=Customer+Stories) # Integrate existing stack If you haven't started building your SaaS Product, you can directly [build](https://docs.omnistrate.com/getting-started/overview/index.md) your SaaS control plane with Omnistrate. However, if you have already built your initial SaaS Product, you can use Omnistrate to extend your current offering to offer even more value. There are a few common ways that our customers extend their current offerings and we will discuss each of those best practices here. ## Extending your SaaS to other usecases (vertical integration) The old and traditional way is to continue to hire engineers and reinvent the wheel, or you can switch to Omnistrate platform and solve all of the undifferentiated problems at a fraction of the cost. Here are some of the possible use-cases that you may want to extend your SaaS: - Go Multi-cloud: You have built the SaaS for AWS or one of the cloud providers and trying to go multi-cloud - Add BYOA support: You have built a SaaS deploying everything in your account but you want to deploy your tech stack in your customers account - Add more Plans: You have built a dedicated offering but want to extend it to other offerings (Multi-Tenant or Serverless) - Go Global: You have built your SaaS in one or two regions but you want to now scale to 130+ regions across major cloud providers or good subset of them - Add new services: You want to extend your SaaS with new product offerings but the cost to hire or extend your current SaaS is too much - Operational cost: You have hired 100's of engineers to operate your SaaS at scale and you want to improve your margins by 10x by moving your existing offering Now, you may have already built something and want to keep it running for your existing customers. Well, one of the architectural pattern that many of our customers used is SaaS gateway pattern. Basically, in this pattern, you keep your UI and the API interface as it is, and simply call Omnistrate generated APIs to add support for clouds/regions/BYOA/serverless etc for your SaaS. Here is a pictorial view of some of the above options: That's great but you may be wondering, how will I integrate: - Authentication: we suggest that you continue to use your existing authentication mechanism and so don't have to worry about this layer at all. - Authorization: if you already have an access control, you can continue to use your access control and only call your Omnistrate-based SaaS APIs for all the deployment and management operations, when the access is granted. - Auditing: we will audit all of your control plane operations. You can call our APIs to get a list of notifications for any tenant or resources, and append it to your existing auditing mechanism. - Metrics/Logging: we integrate with common Observability vendors and allow you to continue to leverage them without any extra effort. - Alerting: we seamlessly integrate with PagerDuty today and provide you alerting out of the box. - Billing: If you have already built a billing pipeline, we can send you the aggregated metering records for your pipeline to consume. Alternatively, we can take care of end to end billing. - Continuous Integration: You may already have or using a CI system. We can directly integrate with your existing CI system. You can then release your new software to your existing SaaS stack and Omnistrate-based stack at the same time. As an example, one of the companies who already had the cloud and a large team to operate the cloud. They wanted to launch a new service and their estimates were a dozen of engineers for 1 year just to build MVP. We were able to get them live in no time using the above architecture and integrate with their existing offering. According to their SVP of engineering "Partnering with Omnistrate was a pivotal decision for our SaaS project. Their offering not only expedited our development process but also ensured a timely launch, allowing us to capitalize on a crucial market opportunity.". If you have any questions on integration challenges your SaaS control plane, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) and one of the core technical team member would love to discuss it in more detail and share the best practices. ## Extend your SaaS architecture (horizontal integration) In this integration, instead of extending your SaaS to other business use-cases, you can offload parts of SaaS control plane to Omnistrate. There are primarily 5 layers: - Tenant management: this layer is responsible for managing tenants, ex- authentication, authorization, UX - API/CLI/UI, auditing, tenant provisioning and management, tenant isolation - Application management: this layer is responsible for managing application from scaling, patching, to extending application capabilities like multi-zone placement, IP whitelisting etc. - Infra management: this layer is responsible for building and operating infrastructure components, ex- compute, storage, networking, load-balancers, certificates, routing, DNS, K8s, infra-security and policies, cloud-accounts, infra-observability, infra-evolution, infra-upgrades etc. - Commerce management: this layer is responsible for metering and billing your customers. - Operational management: this layer is responsible for the operational management of your customers. ### Custom Tenant Management layer If you would like to maintain your own tenant-management layer, you can ignore all the tenant-management APIs that we generated and just use your control plane APIs to build/operate service/infra layers. ### Custom Application Management layer One of the typical mechanism to build application management layer is using the K8s operator framework. Let's say you already have an K8s Operator that defines setup, perform upgrades, take backups, manage scaling etc. and you want to use your Operator to build your SaaS. You can offload the other layers to us by bringing your own operator and we will skip application management. To build and operate your SaaS, you will still need the other SaaS layers beyond K8s operator to manage infrastructure, add SaaS capabilities, enable SaaS integrations, automate day-2 operations etc. As an example, one of the companies already had an operator and wanted to continue to use that. They wanted to launch their SaaS Product and decided to use Omnistrate for everything else. For a concrete operator example, see the [Operator-powered SaaS use case](https://docs.omnistrate.com/usecases/operator-powered-saas/index.md) and the [Build with Kubernetes Operators guide](https://docs.omnistrate.com/build-guides/operators/index.md). The pattern keeps your operator and CRDs as the application lifecycle layer, uses Argo Workflow-style `systemWorkflows` and `customWorkflows` for platform operations, and lets Omnistrate provide the Customer Portal, APIs, tenant management, deployment cells, subscriptions, backups, restores, and fleet operations around it. ### Custom Infrastructure Management layer If you already have Terraform or OpenTofu scripts (or other setup) and prefer to maintain that yourself, you can offload the other layers to us and we will skip creating/managing infrastructure. ### Custom Commerce Management layer You can integrate with billing provider of your choice and don't have to use any of our integrations. To learn more, please see [here](https://docs.omnistrate.com/fin-ops-guides/billing/#byo-billing-provider-bring-your-own-billing-provider) ### Custom Operational Management layer If you would like to automate your day-2 operations outside Omnistrate, you can simply build and manage your SaaS Product with Omnistrate. If you have any questions on integration challenges your SaaS control plane, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) and one of the core technical team member would love to discuss it in more detail and share the best practices. # Internal SaaS / PaaS Internal SaaS refers to a model where software applications are hosted, and maintained by an organization for its own use or for a specific group of users. On the one side is the Platform or IT teams offering those applications running on your servers, and on the other side are the users in rest of the organization. Here is how it may look like: ## Internal SaaS / PaaS benefits Omnistrate enables platform or IT teams to build your Internal SaaS to streamline hosting and management of internal applications to rest of the organization so that teams can focus on things that matters the most for your business from defining policies to exploring new technologies to helping applications teams with their deployment or operational or infrastructure designs. Secondly, standardizing the tooling will enable you to quickly offer a new application to the organization in no time. It will allow you to run experiements, adopt new technologies, deprecate old ones without much overhead. Thirdly, standardizing will streamline the processes and further reduce the cost to operate beyond all the automation offered by the internal saas design. It will allow you to unify policies across the teams, track costs across the teams, rollout security patches quickly, isolate different teams tech-stack as needed, apply operational improvements in one place. Finally, it improves the user-experience as it brings consistency for your users, makes it easy for them to onboard a new team member without having to learn hundreds of different ways to operate. ## Internal SaaS / PaaS features Here are some of the things that your Internal SaaS can help: - Create service templates to enable self-service deployments. Here are some example use-cases: - Deploy across accounts: as you scale, you may want to go across multiple accounts for DR reasons and quota limitations - Go multi-cloud: you may want to keep the flexibility to go across clouds for price negotiation and DR reasons - Go Global: you may want to go global and have requirements like GDPR to keep the data local - Deploy across environments: you may want to have dev, staging and prod environments and keep them in sync - Enable on-demand customization: you may have different configurations in which you want to test your stack - Create self-service portals and empower your application teams to innovate faster - Ability to auto-pause and auto-scale your workloads - Auto-tag your stacks to gain visibility into costs - Ability to update the software for tenants or isolate different tenants - Multi account management and consolidation - Automate day-2 operations from rotating certificates, in-depth observability, upgrading infrastructure, to recovering from different failure scenarios - Automated infrastructure management - Simplify your cloud migration Omnistrate also integrates with your favorite tooling and you can integrate it further with other tools of your choice. You can read more about the supported integrations [here](https://docs.omnistrate.com/build-guides/integrations/index.md) ## Reference architecture for Internal SaaS / PaaS Here is a reference architecture that your Internal SaaS might look like: ## Internal SaaS / PaaS example One of our customers have an Nginx based application with Redis and PostgreSQL components. To keep things simple, we will a server microservice with workers to distribute the work. They wanted to create different stacks in every region for isolation and compliance, have different testing environments and keep it in-sync with the production environment, different configuration to test with different settings. They were able to model their application stack to create a Internal SaaS Product and power their application teams. # Open source SaaS / PaaS Omnistrate offers a comprehensive control plane generator for open-source projects seeking to monetize their offerings and transform them into fully-fledged SaaS products. With Omnistrate, open-source projects can generate private control planes with advanced infrastructure, monetization, and operational capabilities to build, deploy, and manage their SaaS applications effectively. As mentioned [here](https://a16z.com/open-source-from-community-to-commercialization/), open-source company has to go through the 3-phases: 1. Project-community fit, where your open source project creates a community of developers who actively contribute to the open source code base. This can be measured by GitHub stars, commits, pull requests or contributor growth. 1. Product-market fit, where your open source software is adopted by users. This is measured by downloads and usage. 1. Value-market fit, where you find a value proposition that customers want to pay for. The success here is measured by revenue. Omnistrate aims to simplify the last and often the most difficult leg of the journey by automating the undifferentiated components and enabling you to build your SaaS in no time with enterprise grade capabilities at one-tenth of the price. ## Open source SaaS / PaaS benefits A typical startup journey looks like this: you have an idea, you validate the market, you build your initial MVP, you raise some capital. After raising the initial capital or in some cases just before, you are looking to monetize by finding a product-market fit. When we build an application, we typically don’t think about building an operating system or database as it makes no sense to spend time and energy and instead just use the existing components. In the same way, Omnistrate generates the SaaS control plane layer so you can SaaSify your application software without spending the majority of your engineering resources building and operating that layer yourself. Here are some of the key benefits: - **Time to market**: Building control planes can take years to build. Here is an excellent blog from ClickHouse on how it took their world-class team a year to launch their initial SaaS (in AWS only in a subset of the regions). Let’s be honest; not all of us can afford to raise anywhere close to $250MM at a $2B valuation immediately after founding the company. Moreover, not all of us have the luxury to hire an industry veteran in this space like Yury Izrailevsky (ex-VP of Engineering at Google) to lead the team. - **10x Cost effective**: Building control planes are expensive and typical successful SaaS startup run hundreds of engineers to build and maintain just the control plane components. To learn more, please see [here](https://blog.omnistrate.com/posts/52) - **Focus on your differentiation and core strength**: Making startups successful are extremely hard. By streamlining your businesses to stay agile and dedicated to what sets them apart in the market, you can move fast and get the enterprise-grade capabilities at a fraction of price. Not only it will you apart from your competitors but also protect you from the risk of cloud provider picking your technology and start hosting it themselves. To learn more on what we offer out of box, please see [this link](https://docs.omnistrate.com/what-is-omnistrate/index.md) # Operator-powered SaaS Many teams already have a Kubernetes Operator that knows how to run their product: it creates clusters, rolls upgrades, manages backup resources, or reconciles product-specific custom resources. The operator solves the application lifecycle problem, but it does not by itself create a SaaS control plane. To launch a customer-facing product, you still need: - Tenant onboarding and subscription management. - A Customer Portal and generated APIs. - Cloud account onboarding and deployment-cell management. - Per-tenant provisioning, scaling, stopping, starting, backups, restores, and deletes. - RBAC, audit logs, workflow visibility, support tooling, and fleet operations. - Metering, billing, alerts, observability, and upgrade workflows. Omnistrate lets you keep your operator and build the rest of the product around it. ## What Problem This Solves Without Omnistrate, teams often need to build a custom control plane around their operator. That means writing APIs, forms, tenant models, subscription approval flows, deployment orchestration, backup and restore orchestration, status polling, logging, operational dashboards, and cloud-account management. With Omnistrate, your operator stays responsible for reconciling Kubernetes resources, while Omnistrate provides the product control plane and lifecycle orchestration. You define a service plan once and Omnistrate exposes it through the Customer Portal, APIs, CLI, and Operations Center. ## Reference Example The [operator spec template](https://github.com/omnistrate-community/operator-spec-template) shows a complete CloudNativePG PostgreSQL service plan: - CloudNativePG is installed as an operator Helm dependency. - A customer provisions a PostgreSQL cluster from the Omnistrate Customer Portal or API. - Omnistrate passes customer inputs such as instance type, storage size, database name, and replica count into the workflow context. - Required `systemWorkflows.create`, `modify`, and `delete` hooks provision, update, and remove the CNPG `Cluster` custom resource. - `systemWorkflows.stop` and `start` toggle CNPG hibernation. - `systemWorkflows.backup` creates a CNPG `Backup` resource and captures provider backup metadata. - `systemWorkflows.restore` creates a new target cluster from the selected Omnistrate snapshot and the captured backup metadata. - `systemWorkflows.deleteBackup` cleans up provider-side backup resources. This pattern applies to other operators as well: databases, queues, search engines, AI platforms, data pipelines, or internal platforms that already have a CRD-driven lifecycle. If your product needs operations beyond the standard platform lifecycle APIs, you can add provider-defined `customWorkflows` separately. ## No Lock-In to a Proprietary Lifecycle Model Omnistrate is designed to work with the artifacts your platform team already uses: - **Kubernetes Operators and CRDs** remain the source of truth for application reconciliation. - **Helm charts** can install the operator and CRDs. - **Argo Workflow-style YAML** defines lifecycle workflows using DAGs, parameters, and Kubernetes resource actions. - **Terraform or OpenTofu** can optionally manage cloud resources such as buckets, IAM roles, KMS keys, VPC attachments, or external services. - **Your existing portal or API** can call Omnistrate APIs if you do not want to use the generated Customer Portal directly. The result is an end-to-end SaaS solution without rewriting your operator or moving product lifecycle logic into a proprietary format. You keep standard Kubernetes and infrastructure-as-code assets, and Omnistrate adds the control plane surfaces around them. ## End-to-End Architecture A common operator-powered SaaS architecture looks like this: 1. **Service plan definition** You define a plan spec with customer inputs, endpoint configuration, operator dependencies, backup policy, and system workflows. Provider-defined custom workflows can be added when the product needs operations outside the standard lifecycle APIs. 1. **Customer onboarding** Customers sign up through the generated Customer Portal, your custom portal, or your API integration. Omnistrate manages tenant records, subscriptions, permissions, and deployment requests. 1. **Provisioning** Omnistrate creates or selects a deployment cell, creates the tenant namespace, installs required dependencies, and runs the `create` workflow. The workflow applies operator-managed custom resources. 1. **Day-2 operations** Customers or operators can scale, stop, start, back up, restore, delete, or invoke optional provider-defined operations. Omnistrate runs the corresponding Argo Workflow-style lifecycle definition and records debug events for workflow steps. 1. **Optional cloud infrastructure** If the operator needs cloud resources, Terraform can be part of the same plan. For example, Terraform can create an S3 bucket or IAM role, and the operator workflow can use those outputs when creating backup resources. 1. **Operations and support** Your operators use the Operations Center to inspect workflows, debug events, deployment cells, snapshots, alerts, logs, and fleet health. ## When to Use This Pattern Use this pattern when: - You already have an operator and want to commercialize it as a SaaS Product. - Your product requires lifecycle operations beyond create and delete. - You need customer-facing backup and restore backed by operator-native backup resources. - You want to expose controlled day-2 actions such as diagnostics or other provider-defined administrative operations. - You need to support hosted, BYOC, private networking, or more than one cloud/region without building each control plane flow yourself. Start with [Build from Kubernetes Operators](https://docs.omnistrate.com/getting-started/build-from-operators/index.md), then use the full [Build with Kubernetes Operators](https://docs.omnistrate.com/build-guides/operators/index.md) guide when you are ready to model workflows and backup/restore behavior. # Use Cases This section explores various deployment models and use cases for building SaaS Products with Omnistrate. Each use case addresses specific business requirements and deployment scenarios that organizations commonly face when building and scaling their Plans. For more detailed technical information about building SaaS Products, see our [Build Guide](https://docs.omnistrate.com/build-guides/api-params/index.md) section. ## Build your Hosted SaaS Product ### [PaaS](https://docs.omnistrate.com/usecases/paas/index.md) In this Hosted SaaS model, you host the infrastructure in your account and your customers can simply self-serve themselves to deploy the product i.e. they can request a deployment on demand, configure it to their needs, and start using it immediately. In other words, this model enables Product-Led Growth (PLG) by removing manual bottlenecks from provisioning. Examples: AWS RDS, Confluent Cloud (Dedicated) ### [SaaS](https://docs.omnistrate.com/usecases/saas/index.md) In this Hosted SaaS model, your customers don’t even create or configure deployments. The infrastructure is hosted within your account. They simply go to an URL, authenticate themselves and start using your application — no need to create/customize/govern/manage deployments. The endpoint is smart enough to handle a broad range of use cases and automatically adjust based on usage to optimize for the cost, scale, performance without compromising on the security or reliability. Examples: Confluent Cloud (Standard/Enterprise), HubSpot, Rippling. ### [Agent-aaS](https://docs.omnistrate.com/usecases/agent-aas/index.md) Traditional apps are static — users must manually navigate interfaces, trigger workflows, and interpret dashboards. Instead, agents can execute tasks on behalf of users — interacting with APIs, orchestrating services, and adapting based on environment or intent. To build dynamic workflows, agents are not optional — they are essential. But here's the challenge: Agents aren’t monolithic binaries or static microservices. They are iterative, context-driven, often multi-step systems. Hence, we need a separate distribution model for you to distribute them natively. ## Deploy in your Customer Accounts ### [BYOC (Bring Your Own Cloud)](https://docs.omnistrate.com/usecases/byoc/index.md) Deploy your SaaS Products directly in your customers' cloud accounts for enhanced security, compliance, and cost control. This model is ideal for enterprise customers who prefer to maintain data sovereignty and control over their infrastructure while benefiting from your Plans. Examples: Databricks Cloud, StarTree BYOC, Redpanda BYOC. ### [BYOC On-Premise](https://docs.omnistrate.com/usecases/byoc-onprem/index.md) Run your product inside a customer-owned Kubernetes cluster while retaining lifecycle operations, observability, and upgrade workflows through your control plane. ### [Air-Gapped](https://docs.omnistrate.com/usecases/air-gapped/index.md) Enable deployments with no external dependencies while maintaining centralized management capabilities. Perfect for customers with strict security requirements or air-gapped environments who need enterprise-grade functionality within their own infrastructure. ## Other use cases ### [Integrate Existing Stack](https://docs.omnistrate.com/usecases/integrate-existing-stack/index.md) Extend your current platform with new distribution channels without rebuilding from scratch. Learn about vertical integration patterns (adding new deployment models, regions, or services) and horizontal integration patterns (offloading specific layers of your SaaS Product to Omnistrate). ### [Internal SaaS/PaaS](https://docs.omnistrate.com/usecases/internal-saas/index.md) Build internal platforms to streamline application hosting and management within your organization. Enable platform and IT teams to offer standardized services to internal teams, improving efficiency, reducing operational overhead, and maintaining consistency across deployments. ### [Open Source SaaS/PaaS](https://docs.omnistrate.com/usecases/open-source/index.md) Transform open-source projects into monetizable Plans with enterprise-grade capabilities. Accelerate the journey from open-source community adoption to commercial success by leveraging Omnistrate's platform to handle the undifferentiated heavy lifting of operations. ### [Operator-powered SaaS](https://docs.omnistrate.com/usecases/operator-powered-saas/index.md) Bring an existing Kubernetes Operator and build the SaaS control plane around it. Use standard CRDs, Helm, Argo Workflow-style lifecycle definitions, optional Terraform, and Omnistrate's generated Customer Portal, APIs, tenant management, backups, restores, and fleet operations. # PaaS (Platform as a Service) ## What is PaaS? Platform as a Service (PaaS) is a cloud computing model that provides a complete development and deployment environment in the cloud. PaaS sits between Infrastructure as a Service (IaaS) and Software as a Service (SaaS) in the cloud service stack, offering a middle layer that abstracts away infrastructure complexity while providing the tools and services needed to build, deploy, and manage applications. Unlike OnPrem deployments, PaaS enables your clients to self-serve and deploy your product on demand. They can request a deployment, configure it to their needs, and start using it immediately. This model enables Product-Led Growth (PLG) by removing manual bottlenecks from provisioning. Examples include AWS RDS and Confluent Cloud (Dedicated). PaaS is ideal for SMBs and enterprises that have found product-market fit and want to scale their PLG motion by building self-serve deployments, usage-based billing, and automated operations at massive scale — reliably, securely, and cost-effectively. ### Common PaaS Challenges - **Seamless tenant onboarding and self-serve deployments with guardrails** - **End-to-End usage-based metering and billing** - **Automated operations to be able to operate at scale** - **Supporting a wide range of customer journeys** — from trials to enterprise-grade deployments without rebuilding distribution infrastructure at every stage ### Key Characteristics of PaaS **For Service Providers:** - **Abstraction Layer**: Hide infrastructure complexity from customers - **Managed Services**: Handle provisioning, scaling, monitoring, and maintenance automatically - **API-Driven**: Provide programmatic access to platform capabilities - **Self-Service**: Enable customers to provision and manage resources independently **For Customers:** - **No Infrastructure Management**: Focus on application development rather than server management - **Rapid Deployment**: Deploy applications quickly without setting up underlying infrastructure - **Scalability**: Automatically scale resources based on demand - **Cost Efficiency**: Pay only for resources used, with no upfront infrastructure costs - **Built-in Services**: Access databases, monitoring, security, and other services out of the box ### Common PaaS Use Cases - **Database as a Service (DBaaS)**: Managed PostgreSQL, MySQL, MongoDB, Redis - **Application Hosting**: Runtime environments for web applications and APIs - **Analytics Platforms**: Data processing, business intelligence, machine learning - **Integration Platforms**: API management, message queues, event streaming ## Building PaaS with Omnistrate Platform as a Service (PaaS) enables you to offer managed database and application services to your customers without them having to worry about the underlying infrastructure. With Omnistrate, you can quickly transform your database or application into a fully managed PaaS offering that your customers can consume on-demand. PaaS offerings typically include: - **Managed Databases**: PostgreSQL, MySQL, MongoDB, Redis, and other data stores - **Application Platforms**: Runtime environments for various programming languages - **Development Tools**: CI/CD pipelines, testing environments, and deployment tools - **Middleware Services**: Message queues, caching layers, and API gateways ## How Omnistrate Enables PaaS Success Building and operating a successful PaaS requires expertise across three critical areas: **Operations**, **Distribution**, and **Monetization**. Omnistrate provides comprehensive solutions for each: ### Operations Excellence **Automated Infrastructure Management** - **Multi-Cloud Orchestration**: Deploy and manage services across AWS, GCP, and Azure from a single platform - **Auto-Scaling**: Automatically scale resources based on demand with configurable thresholds - **Self-Healing**: Detect and recover from failures automatically with built-in health checks - **Zero-Downtime Updates**: Roll out updates and patches without service interruption - **Backup & Recovery**: Automated backup scheduling with point-in-time recovery capabilities **Day-2 Operations** - **Monitoring & Alerting**: Comprehensive observability with metrics, logs, and custom alerts - **Security Management**: Automated security patching, certificate rotation, and compliance monitoring - **Performance Optimization**: Continuous performance tuning and resource optimization - **Multi-Tenancy**: Secure isolation between customer instances with shared infrastructure efficiency - **Disaster Recovery**: Cross-region replication and automated failover capabilities **Operational Efficiency** - **Infrastructure as Code**: Version-controlled infrastructure definitions with GitOps workflows - **Policy Enforcement**: Automated governance and compliance across all deployments - **Cost Optimization**: Intelligent resource allocation and usage-based cost tracking - **Audit Trails**: Complete audit logs for compliance and troubleshooting ### Distribution & Go-to-Market **Multi-Channel Distribution** - **Self-Service Portals**: White-labeled customer portals for service provisioning and management - **API-First Architecture**: RESTful APIs for programmatic access and third-party integrations - **Marketplace Integration**: Native integration with AWS, GCP, and Azure marketplaces - **Partner Channels**: Enable reseller and partner distribution models - **Developer Experience**: SDKs, CLI tools, and comprehensive documentation **Global Reach** - **Multi-Region Deployment**: Deploy services in multiple regions for low latency and compliance - **BYOC Support**: Enable customers to deploy in their own cloud accounts for security and compliance - **Hybrid Deployment**: Support both Hosted SaaS and BYOC deployment models - **Edge Computing**: Deploy services closer to end-users for improved performance **Customer Onboarding** - **Instant Provisioning**: Deploy new instances in minutes, not hours or days - **Trial Management**: Automated trial provisioning with usage limits and expiration - **Migration Tools**: Assist customers in migrating from existing solutions - **Documentation & Support**: Auto-generated documentation and integrated support systems ### Monetization & Business Growth **Flexible Pricing Models** - **Usage-Based Billing**: Charge customers based on actual resource consumption - **Subscription Tiers**: Create multiple service tiers with different feature sets - **Marketplace Billing**: Leverage cloud marketplace billing for simplified customer acquisition - **Custom Pricing**: Support enterprise contracts with custom pricing structures - **Freemium Models**: Offer free tiers to drive adoption and upselling **Revenue Operations** - **Automated Metering**: Track resource usage across all customer instances automatically - **Billing Integration**: Connect with popular billing systems like Stripe, Chargebee, and others - **Cost Attribution**: Understand the true cost of serving each customer - **Revenue Analytics**: Detailed insights into revenue trends, customer lifetime value, and churn - **Invoicing & Collections**: Automated invoice generation and payment processing ## PaaS Benefits with Omnistrate Building a PaaS with Omnistrate provides several key advantages: - **Rapid Time-to-Market**: Transform your existing database or application into a managed service in days, not months - **Multi-Cloud Support**: Deploy across AWS, GCP, and Azure without vendor lock-in - **Automated Operations**: Handle provisioning, scaling, backups, monitoring, and maintenance automatically - **Customer Self-Service**: Provide intuitive portals for customers to manage their instances - **Flexible Deployment Models**: Support both Hosted SaaS and BYOC deployments - **Built-in Observability**: Comprehensive metrics, logging, and alerting out of the box - **Scalable Business Model**: Start small and scale to enterprise with built-in monetization tools - **Reduced Operational Overhead**: Focus on your core product while Omnistrate handles infrastructure complexity ## PostgreSQL PaaS Example Let's walk through creating a PostgreSQL PaaS service using Omnistrate. We'll start with a basic setup and incrementally add enterprise features. ### Basic PostgreSQL Service Here's a simple PostgreSQL service definition: ``` version: '3.9' services: PostgreSQL: image: bitnami/postgresql:latest ports: - 5432:5432 volumes: - ./data:/var/lib/postgresql/data environment: - POSTGRESQL_PASSWORD=defaultpassword - POSTGRESQL_DATABASE=postgres - POSTGRESQL_USERNAME=postgres - POSTGRESQL_POSTGRES_PASSWORD=your-secure-password ``` While this works, it lacks the enterprise features needed for a production PaaS offering. Let's enhance it step by step. ### Step 1: Configure Cloud Provider Account First, specify where to deploy your PostgreSQL instances. You can deploy in your own account (Hosted SaaS) or your customers' accounts (BYOC). #### Hosted SaaS Deployment (Multi-Cloud) ``` x-omnistrate-service-plan: name: 'PostgreSQL Service' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' ``` #### BYOC (Bring Your Own Cloud) Support ``` byoaDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role ``` ### Step 2: Add Observability and Integrations Enable comprehensive monitoring and logging: ``` x-customer-integrations: logs: metrics: ``` ### Step 3: Add Customer Customization Allow customers to configure their PostgreSQL instances: ``` services: PostgreSQL: image: bitnami/postgresql:latest ports: - 5432:5432 volumes: - ./data:/var/lib/postgresql/data x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: instanceType - cloudProvider: gcp apiParam: instanceType - cloudProvider: azure apiParam: instanceType x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP3 instanceStorageSizeGiAPIParam: storageSize instanceStorageIOPSAPIParam: storageIOPS gcp: instanceStorageType: GCP::PD_SSD instanceStorageSizeGiAPIParam: storageSize azure: instanceStorageType: AZURE::PREMIUM_LRS instanceStorageSizeGiAPIParam: storageSize environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_DATABASE=$var.postgresqlDatabase - POSTGRESQL_USERNAME=$var.postgresqlUsername - POSTGRESQL_POSTGRES_PASSWORD=$var.postgresqlRootPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true x-omnistrate-api-params: - key: instanceType description: Instance Type for PostgreSQL name: Instance Type type: String modifiable: true required: true export: true - key: storageSize description: Storage Size in GB name: Storage Size (GB) type: Float64 modifiable: true required: true export: true defaultValue: "100" limits: min: 20 max: 1000 - key: storageIOPS description: Storage IOPS name: Storage IOPS type: Float64 modifiable: true required: false export: true defaultValue: "3000" - key: postgresqlUsername description: Database Username name: Database Username type: String modifiable: false required: true export: true - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false - key: postgresqlDatabase description: Database Name name: Database Name type: String modifiable: false required: true export: true defaultValue: "postgres" - key: postgresqlRootPassword description: Root Password name: Root Password type: String modifiable: false required: false export: false defaultValue: "generate-secure-password" ``` ### Step 4: Add High Availability with Read Replicas Create a master-replica setup for high availability: ``` services: Master: image: bitnami/postgresql:latest ports: - 5432:5432 volumes: - ./data:/var/lib/postgresql/data x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: masterInstanceType - cloudProvider: gcp apiParam: masterInstanceType - cloudProvider: azure apiParam: masterInstanceType x-omnistrate-capabilities: enableEndpointPerReplica: true enableMultiZone: true environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_DATABASE=$var.postgresqlDatabase - POSTGRESQL_USERNAME=$var.postgresqlUsername - POSTGRESQL_POSTGRES_PASSWORD=$var.postgresqlRootPassword - POSTGRESQL_REPLICATION_MODE=master - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=$var.replicationPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true x-omnistrate-api-params: - key: masterInstanceType description: Master Instance Type name: Master Instance Type type: String modifiable: true required: true export: true - key: postgresqlUsername description: Database Username name: Database Username type: String modifiable: false required: true export: true - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false - key: postgresqlDatabase description: Database Name name: Database Name type: String modifiable: false required: true export: true - key: postgresqlRootPassword description: Root Password name: Root Password type: String modifiable: false required: false export: false defaultValue: "rootpassword123" - key: replicationPassword description: Replication Password name: Replication Password type: String modifiable: false required: false export: false Replica: image: bitnami/postgresql:latest ports: - 5433:5432 volumes: - ./data:/var/lib/postgresql/data x-omnistrate-compute: replicaCountAPIParam: numReadReplicas instanceTypes: - cloudProvider: aws apiParam: replicaInstanceType - cloudProvider: gcp apiParam: replicaInstanceType - cloudProvider: azure apiParam: replicaInstanceType x-omnistrate-capabilities: enableMultiZone: true enableEndpointPerReplica: true environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_MASTER_HOST=Master - POSTGRESQL_MASTER_PORT_NUMBER=5432 - POSTGRESQL_REPLICATION_MODE=slave - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=$var.replicationPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true x-omnistrate-api-params: - key: replicaInstanceType description: Replica Instance Type name: Replica Instance Type type: String modifiable: true required: true export: true - key: numReadReplicas description: Number of Read Replicas name: Number of Read Replicas type: Float64 modifiable: true required: false export: true defaultValue: "1" limits: min: 0 max: 5 - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false - key: replicationPassword description: Replication Password name: Replication Password type: String modifiable: false required: false export: false defaultValue: "replpassword123" ``` ### Step 5: Add Enterprise Features Enable autoscaling, backup, and advanced capabilities: ``` x-omnistrate-capabilities: enableMultiZone: true enableEndpointPerReplica: true autoscaling: minReplicas: 1 maxReplicas: 10 idleMinutesBeforeScalingDown: 5 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 2 overUtilizedThreshold: 80 backupConfiguration: backupRetentionInDays: 7 backupPeriodInHours: 2 ``` ### Complete PostgreSQL PaaS Example Here's the complete compose specification for a production-ready PostgreSQL PaaS: ``` version: '3.9' x-omnistrate-service-plan: name: 'PostgreSQL Service' tenancyType: 'OMNISTRATE_DEDICATED_TENANCY' deployment: hostedDeployment: awsAccountId: "" awsBootstrapRoleAccountArn: arn:aws:iam:::role/omnistrate-bootstrap-role gcpProjectId: "" gcpProjectNumber: "" gcpServiceAccountEmail: "" azureSubscriptionId: '' azureTenantId: '' x-customer-integrations: logs: metrics: services: Master: image: bitnami/postgresql:latest ports: - 5432:5432 volumes: - ./data:/var/lib/postgresql/data x-omnistrate-compute: instanceTypes: - cloudProvider: aws apiParam: masterInstanceType - cloudProvider: gcp apiParam: masterInstanceType - cloudProvider: azure apiParam: masterInstanceType x-omnistrate-storage: aws: instanceStorageType: AWS::EBS_GP3 instanceStorageSizeGiAPIParam: storageSize instanceStorageIOPSAPIParam: storageIOPS gcp: instanceStorageType: GCP::PD_SSD instanceStorageSizeGiAPIParam: storageSize azure: instanceStorageType: AZURE::PREMIUM_LRS instanceStorageSizeGiAPIParam: storageSize x-omnistrate-capabilities: enableEndpointPerReplica: true enableMultiZone: true autoscaling: minReplicas: 1 maxReplicas: 3 idleMinutesBeforeScalingDown: 5 idleThreshold: 20 overUtilizedMinutesBeforeScalingUp: 2 overUtilizedThreshold: 80 environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_DATABASE=$var.postgresqlDatabase - POSTGRESQL_USERNAME=$var.postgresqlUsername - POSTGRESQL_POSTGRES_PASSWORD=$var.postgresqlRootPassword - POSTGRESQL_REPLICATION_MODE=master - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=$var.replicationPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true - POSTGRESQL_DATA_DIR=/var/lib/postgresql/data/dbdata - SECURITY_CONTEXT_USER_ID=1001 - SECURITY_CONTEXT_FS_GROUP=1001 - SECURITY_CONTEXT_GROUP_ID=0 x-omnistrate-api-params: - key: masterInstanceType description: Master Instance Type name: Master Instance Type type: String modifiable: true required: true export: true - key: storageSize description: Storage Size in GB name: Storage Size (GB) type: Float64 modifiable: true required: true export: true defaultValue: "100" limits: min: 20 max: 1000 - key: storageIOPS description: Storage IOPS name: Storage IOPS type: Float64 modifiable: true required: false export: true defaultValue: "3000" - key: postgresqlUsername description: Database Username name: Database Username type: String modifiable: false required: true export: true - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false - key: postgresqlDatabase description: Database Name name: Database Name type: String modifiable: false required: true export: true defaultValue: "postgres" - key: postgresqlRootPassword description: Root Password name: Root Password type: String modifiable: false required: false export: false defaultValue: "CHANGEME_SECURE_PASSWORD" - key: replicationPassword description: Replication Password name: Replication Password type: String modifiable: false required: false export: false defaultValue: "CHANGEME_SECURE_PASSWORD" x-omnistrate-mode-internal: true Replica: image: bitnami/postgresql:latest ports: - 5433:5432 volumes: - ./data:/var/lib/postgresql/data x-omnistrate-compute: replicaCountAPIParam: numReadReplicas instanceTypes: - cloudProvider: aws apiParam: replicaInstanceType - cloudProvider: gcp apiParam: replicaInstanceType - cloudProvider: azure apiParam: replicaInstanceType x-omnistrate-capabilities: enableMultiZone: true enableEndpointPerReplica: true environment: - POSTGRESQL_PASSWORD=$var.postgresqlPassword - POSTGRESQL_MASTER_HOST=Master - POSTGRESQL_MASTER_PORT_NUMBER=5432 - POSTGRESQL_REPLICATION_MODE=slave - POSTGRESQL_REPLICATION_USER=repl_user - POSTGRESQL_REPLICATION_PASSWORD=$var.replicationPassword - POSTGRESQL_PGAUDIT_LOG=READ,WRITE - POSTGRESQL_LOG_HOSTNAME=true - POSTGRESQL_DATA_DIR=/var/lib/postgresql/data/dbdata - SECURITY_CONTEXT_USER_ID=1001 - SECURITY_CONTEXT_FS_GROUP=1001 - SECURITY_CONTEXT_GROUP_ID=0 x-omnistrate-api-params: - key: replicaInstanceType description: Replica Instance Type name: Replica Instance Type type: String modifiable: true required: true export: true - key: numReadReplicas description: Number of Read Replicas name: Number of Read Replicas type: Float64 modifiable: true required: false export: true defaultValue: "1" limits: min: 0 max: 5 - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false - key: replicationPassword description: Replication Password name: Replication Password type: String modifiable: false required: false export: false defaultValue: "CHANGEME_SECURE_PASSWORD" x-omnistrate-mode-internal: true Cluster: image: omnistrate/noop x-omnistrate-api-params: - key: instanceType description: Instance Type name: Instance Type type: String modifiable: true required: true export: true parameterDependencyMap: Master: masterInstanceType Replica: replicaInstanceType - key: postgresqlPassword description: Database Password name: Database Password type: String modifiable: false required: true export: false parameterDependencyMap: Master: postgresqlPassword Replica: postgresqlPassword - key: postgresqlUsername description: Database Username name: Database Username type: String modifiable: false required: true export: true parameterDependencyMap: Master: postgresqlUsername - key: postgresqlDatabase description: Database Name name: Database Name type: String modifiable: false required: true export: true defaultValue: "postgres" parameterDependencyMap: Master: postgresqlDatabase - key: numReadReplicas description: Number of Read Replicas name: Number of Read Replicas type: Float64 modifiable: true required: false export: true defaultValue: "1" limits: min: 0 max: 5 parameterDependencyMap: Replica: numReadReplicas - key: storageSize description: Storage Size in GB name: Storage Size (GB) type: Float64 modifiable: true required: true export: true defaultValue: "100" limits: min: 20 max: 1000 parameterDependencyMap: Master: storageSize depends_on: - Master - Replica x-omnistrate-mode-internal: false ``` ## Key Features of This PostgreSQL PaaS This example demonstrates several important PaaS capabilities: ### Multi-Cloud Support - Deploy on AWS, GCP, or Azure - Consistent experience across cloud providers - Cloud-specific optimizations (storage types, instance types) ### High Availability - Master-replica architecture - Multi-zone deployment - Automatic failover capabilities ### Customer Customization - Configurable instance types - Adjustable storage size and IOPS - Scalable read replicas (0-5 replicas) ### Enterprise Features - Comprehensive monitoring and logging - Automated backups - Autoscaling based on utilization - Security and compliance best practices ### Deployment Flexibility - Multi-cloud Hosted SaaS option with private networking capabilities - BYOC option for enterprises ## Getting Started with your PaaS To create your PaaS: 1. **Set up your accounts**: Configure your cloud provider accounts using the bootstrap roles 1. **Customize the compose spec**: Modify the example above to match your requirements 1. **Deploy your service**: Use the Omnistrate platform to deploy your PaaS 1. **Test and iterate**: Create test instances and refine your offering 1. **Launch to customers**: Make your PaaS available to customers For more detailed examples and advanced configurations, check out our [examples](https://github.com/omnistrate-community/examples/blob/main/docs/README.md) and [build guides](https://docs.omnistrate.com/build-guides/overview/index.md). For concrete examples of creating a PostgreSQL PaaS using Docker Compose, Helm, and Kubernetes Operators, see the [Omnistrate Community PostgreSQL PaaS repository](https://github.com/omnistrate-community/postgres-paas/). # Building a SaaS Product with Omnistrate Software-as-a-Service (SaaS) has become the dominant model for software delivery, offering customers on-demand access to applications without the complexity of managing underlying infrastructure. In a SaaS model, the provider hosts and manages the application, delivering it to customers over the internet. This approach enables seamless access, automatic updates, and scalable, usage-based pricing. ## What is the SaaS Model? In the Hosted SaaS model, your customers don’t create or configure deployments. The infrastructure is hosted within your account, and users simply access the application through a URL, authenticate themselves, and start using it. The application endpoint is designed to handle a broad range of use cases and automatically adjusts based on usage to optimize for cost, scale, and performance without compromising on security or reliability. Key characteristics of the SaaS model include: - **Automated Infrastructure**: The application is hosted and managed by the provider. - **On-Demand Access**: Users access the software via a web browser or API. - **Automated Operations**: The provider handles all updates, maintenance, and scaling. Examples of successful SaaS products include Salesforce, HubSpot, and Rippling. ## How Omnistrate Helps You Build SaaS Building SaaS requires more than hosting the application: you need repeatable tenant onboarding, infrastructure provisioning, lifecycle operations, billing integration, and fleet management. Omnistrate derives those workflows from your product definition so the SaaS control plane is generated from the same source of truth as your deployment configuration. The sections below break down the core control plane capabilities Omnistrate generates for SaaS delivery. ### Infrastructure Management Omnistrate automates the provisioning and management of your service infrastructure across any cloud provider. You can define your infrastructure requirements as simple abstractions, and Omnistrate’s control plane handles the creation and management of all necessary cloud resources, such as virtual machines, storage, DNS endpoints, and networking rules. This ensures that your infrastructure is version-controlled, repeatable, and can be deployed safely. ### Tenant Management A successful SaaS product requires robust tenant management. Omnistrate enables you to build comprehensive tenant management capabilities that handle the full customer lifecycle, from onboarding to ongoing operations. The platform provides flexible tenancy models, automated deployment orchestration, and integrated billing systems that can adapt to your specific business requirements while maintaining enterprise-grade security and compliance at a low operational cost. Learn more about [Tenant Management](https://docs.omnistrate.com/tenant-management/overview/index.md). ### Operations Management Operating a SaaS application at scale demands proactive monitoring, automated recovery, and intelligent operational workflows. Omnistrate allows you to build a centralized Operations Center that provides: - **Proactive Monitoring**: Continuously monitor your entire deployment fleet to predict and prevent issues. - **Automated Recovery**: Automatically handle common failure scenarios, from machine failures to network partitions, to maintain high availability. - **Intelligent Fleet Management**: Manage upgrades, patching, and configuration updates across your entire fleet with automated, safe rollout strategies. - **Operational Visibility**: Gain a unified view of your fleet's health and performance through a single, AI-powered dashboard. For more details, please see [here](https://docs.omnistrate.com/what-is-omnistrate/#operations-management) ### Application Management Omnistrate provides intelligent runtime management to not only package the application but also ensure your application performs optimally and cost-effectively. Key capabilities include: - **Autoscaling**: Automatically scale your application based on custom metrics to handle variable workloads. - **Serverless Scale-to-Zero**: Reduce costs by automatically scaling resources down to zero during periods of inactivity. - **High Availability**: Implement automated backup and recovery mechanisms to ensure data persistence and service reliability. - **Advanced Networking**: Configure complex networking requirements to meet security and performance needs. For more details, please see [here](https://docs.omnistrate.com/what-is-omnistrate/#application-management) ### Billing Management Omnistrate enables you to build comprehensive FinOps and monetization capabilities that provide accurate usage metering, flexible billing models, and seamless payment processing. The platform automatically tracks resource consumption, calculates charges based on your pricing models, and integrates with existing billing systems to streamline revenue operations. For more details, please see [here](https://docs.omnistrate.com/what-is-omnistrate/#financial-operations-and-monetization) ### Deployment management As your SaaS grows, you’ll need the flexibility to create Cells (isolated deployment units) for a variety of scenarios, including: - Disaster recovery and security containment - Blue/green or canary deployments for safe rollouts - Elastic scaling to handle customer demand - Expanding globally across regions - Meeting data residency and compliance requirements - Serving high-value or regulated customers in dedicated environments For more details, please see [here](https://docs.omnistrate.com/build-guides/deployment-cells/index.md) By leveraging Omnistrate, you can accelerate your journey to delivering a market-ready SaaS product, reducing the operational complexity and engineering overhead traditionally associated with building and scaling SaaS. # Miscellaneous # Omnistrate API resources ## Overview Omnistrate supports REST API as a programmatic alternative to the UI for creating and managing Omnistrate projects. It allows you to automate processes and iterate more quickly, and lets you use Omnistrate from your own UI. You can use the API with Omnistrate supported client in Golang, or with your own custom code. The Golang client is supported in Windows, UNIX, and OS X environments. For manual interaction and CI/CD automation we recommend using Omnistrate [Command Line](https://ctl.omnistrate.cloud/omnistrate-ctl/), which simplifies the interaction with Omnistrate APIs. Note To build your own Customer portal using Omnistrate API check the following [guideline](https://docs.omnistrate.com/tenant-management/build-your-own-portal/index.md). ## Reference documentation Omnistrate offers reference documentation available for the following programmatic tools. | Topic | Description | | ------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [**Omnistrate Build REST API**](https://api.omnistrate.cloud/docs/external/) | The Omnistrate REST API provides a programmatic alternative to the UI for building Omnistrate services. | | [**Omnistrate Consumption REST API**](https://api.omnistrate.cloud/docs/external/#tag/resource-instance-api) | The Omnistrate REST API provides a programmatic alternative to the UI for consuming Omnistrate services. | | [**Omnistrate Fleet REST API**](https://api.omnistrate.cloud/docs/internal/) | The Omnistrate REST API provides a programmatic alternative to the UI for operating Omnistrate services. | | [**OpenAPI specification for Build API**](https://api.omnistrate.cloud/2022-09-01-00/openapi.yaml) | Omnistrate customers can reference the OpenAPI specification for the Omnistrate REST API. It helps automate the generation of a client for languages that Omnistrate doesn't offer a client. It assists with design, implementation, and testing integration with Omnistrate's REST API using a variety of automated OpenAPI-compatible tools. ([yaml](https://api.omnistrate.cloud/2022-09-01-00/openapi.yaml) / [json](https://api.omnistrate.cloud/2022-09-01-00/openapi.json)) | | [**OpenAPI specification for Consumption API**](https://api.omnistrate.cloud/2022-09-01-00/openapi.yaml) | Omnistrate customers can reference the OpenAPI specification for the Omnistrate REST API. It helps automate the generation of a client for languages that Omnistrate doesn't offer a client. It assists with design, implementation, and testing integration with Omnistrate's REST API using a variety of automated OpenAPI-compatible tools. ([yaml](https://api.omnistrate.cloud/2022-09-01-00/openapi.yaml) / [json](https://api.omnistrate.cloud/2022-09-01-00/openapi.json)) | | [**OpenAPI specification for Fleet API**](https://api.omnistrate.cloud/2022-09-01-00/fleet/openapi.yaml) | Omnistrate customers can reference the OpenAPI specification for the Omnistrate REST API. It helps automate the generation of a client for languages that Omnistrate doesn't offer a client. It assists with design, implementation, and testing integration with Omnistrate's REST API using a variety of automated OpenAPI-compatible tools. ([yaml](https://api.omnistrate.cloud/2022-09-01-00/fleet/openapi.yaml) / [json](https://api.omnistrate.cloud/2022-09-01-00/fleet/openapi.json)) | | [**Omnistrate Go SDK**](https://github.com/omnistrate-oss/omnistrate-sdk-go/) | The Omnistrate Go SDK provides an API client to interact with Omnistrate APIs from Go applications and services. The Go SDK includes Build, Consumption and Fleet API references. | # Glossary ## Service A Service represents a SaaS Product that you are trying to build. You can create an empty new service or bring your compose specification or choose from one of the existing compose templates to get started. A service object by itself is a container object to store the service metadata. ## Plan A service consists of several Plans that allows you to offer your service in different Deployment Models and Tenancy Types. For more information on Deployment Models and Tenancy Types, please see [here](https://docs.omnistrate.com/build-guides/deployment-models/index.md) and [here](https://docs.omnistrate.com/build-guides/tenancy-types/index.md) ## Resource Resource (also referred to as resource) is a unit of functionality that can be run independently like a microservice. Your SaaS may have 1 or more Resources. For more information, please see [here](https://docs.omnistrate.com/build-guides/resource/index.md). It is a representation of your SaaS, reflecting its health status, compute/node configuration and other metadata. ## Resource Instance A resource instance is a running instance of a Resource. Users can create and run multiple instances running in different environments (like dev, stage, prod) - regions (like us-east-1, us-west-1) and cloud providers. A resource instance is not to be confused with a container or a Kubernetes pod. ## API Parameters API parameters is a way for you to define the external parameters for your customers to configure a Resource. For more information, please see [here](https://docs.omnistrate.com/build-guides/api-params/index.md) ## Images Images are primarily Container Images that can be configured and attached to the service components ## Infrastructure Infrastructure is your compute/network/storage configuration for different Resources. ## Action Hooks Action hooks allows you to customize your control plane by injecting custom code at different phase of the lifecycle of your SaaS operations. For more information, please see [here](https://docs.omnistrate.com/build-guides/actionhooks/index.md) ## Dependencies Dependencies allows you to specify DAG and how different Resources are dependent on each other. For more information, please see [here](https://docs.omnistrate.com/build-guides/dependencies/index.md) ## Integrations Integrations are 1st/3rd party SaaS applications that you want to integrate with for your SaaS. For more information, please see [here](https://docs.omnistrate.com/build-guides/integrations/index.md) ## Environment Environment allows you setup continuous delivery right in the omnistrate platform to test things in a sandbox environment, or promote changes from developer to stage to production environment. ## Upgrades Upgrades refers to the process of applying updates, fixes, or patches to the software and configurations including infrastructure upgrades. For more information, please see [here](https://docs.omnistrate.com/dev-ops-guides/upgrades/index.md) ## Deployments Deployments refer to the process of deploying your resource instances to a specific version of your Plan. It can be triggered by either creating a new resource instance or patching an existing resource instance. ## Live Deployments Live Deployments are the currently running resource instances of your Resources that have been successfully deployed and are operational in a specific environment. These are active deployments that are serving users or are actively maintained. Monitoring live deployments is crucial to ensure uptime, performance, and the correct functioning of your services. ## Releases Releases refer to the process of releasing a new version of your Plan. Releases are tagged with a version number and can include updates, bug fixes, new features, or other changes. You can mark a release as active, preferred, or deprecated to indicate its status and availability for deployment. ## Alerts Alerts are notifications generated by the system to inform you of important events or issues that require your attention. These could be related to deployment issues, security, or other operational concerns. Alerts are often categorized by priority and will expire in a specific time frame given the severity of the issue. ## Open Alerts Open Alerts are the currently active alerts that have not been resolved. These alerts require immediate attention to prevent service disruptions or other issues. Monitoring open alerts is crucial to ensure the health and stability of your services. ## Notifications Notifications are system-generated messages that inform you about the status and events related to your SaaS Product operations. Unlike Alerts, which are triggered by issues or problems that require immediate attention, Notifications provide general information about routine activities such as successful deployments, plan updates, user signups, or scheduled maintenance. Notifications are informational in nature and typically do not indicate problems that need urgent resolution. ## Signups Signups refer to the number of new users that have registered for your service. Tracking signups is important to measure the growth of your user base and the effectiveness of your marketing and sales efforts. ## Workflows Workflows encompass all user actions to interact with the resource instances. This includes creating, updating, deleting, stopping, starting, patching, and other operations performed on the resource instances. # FAQ ## How much flexibility will I have for customization? As we started working on Omnistrate, one of the core tenets was to allow our customers to extend and customize their control plane combined with best practices so that they don't have to reinvent the wheel again and again on this. - Our architecture is completely plugin based such that you can inject custom code at different phases of the control plane operations. - Extend your control plane operations outside of Omnistrate as you have the full infrastructure control. - We are completely based on open standards. - 200+ APIs to customize your SaaS We also have customers in production including public companies who have thoroughly evaluated and tested our architecture. If you have a specific question or concern, we will be happy to jump on a call and help clarify. We have built dozens of successful cloud services across domains and clouds at AWS, Confluent, and Microsoft. We have used our experience to build a general-purpose SaaS control plane platform that agnostic of any application and scales across different SaaS dimensions like clouds, regions, deployment models, tenancy types, environments and so on. ## What does your operating model looks like? Our vision is to remove the barrier for entrepreneurs and enterprises alike to create their own SaaS businesses. Hence, its paramount for us to fully empower you with all the power and flexibility without requiring any proprietary information to make it an easy decision. More specifically: - We don't require any access to your customers OR your data OR your software - No access or secret keys required. We worked on fine grained permissions that you can revoke at any time - We serve you directly and not your customers. Think of us as any other SaaS vendor like Zendesk, PagerDuty or Hubspot but just focused on helping you build and operate your SaaS ## Whats the responsibility model looks like? For control plane, we take the full responsibility all the way from design choices, preventive measures to operational support. To be more specific, our SaaS control plane architecture is resilient: 1/ decoupled from data plane 2/ per region separation avoiding single point of failure. We have inbuilt safe retries for every control plane operation. If we can't recover, we provide full end-to-end support on the control plane issues with enterprise SLA. In addition, all the user activity will be audited and provide access control (RBAC) for your teams to grant fine-grained privileges. For more on this, please visit [this page](https://docs.omnistrate.com/governance-guides/omnistrate-rbac/index.md). For your application, we have monitoring at different levels across process, pod, machine, network/AZ levels. We do basic recovery from restarting the process, replacing infrastructure etc. For more on this, please visit [this page](https://docs.omnistrate.com/operate-guides/monitoring/index.md). If we still can't recover, we will alert you with relevant debugging information, metrics and logs. If you have any other specific questions, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) ## How do I integrate with Omnistrate if I already have some version of a control plane? Please see [this page](https://docs.omnistrate.com/usecases/integrate-existing-stack/index.md) ## Do you have all the integrations I need? We are adding new integrations as quickly as we can. If you need a specific integration that’s not supported yet, you can request it by [contacting us](https://www.omnistrate.com/contact), or simply create it yourself. For the list of supported integrations, please see [this page](https://docs.omnistrate.com/build-guides/integrations/index.md) ## How does the pricing looks like? We have usage-based [pricing](https://www.omnistrate.com/pricing) and you only pay for what you use. [Contact us](https://www.omnistrate.com/contact) for any questions on pricing. In addition, we offer 2-weeks period to evaluate our product in your account or your customers' account at no cost. ## How long does it take to get started? You should be able to sign up and configure your account in a few minutes. If you already have a Docker Compose file, you can simply follow our [Compose extensions](https://docs.omnistrate.com/build-guides/compose-spec/#custom-tags) to complete your Compose specification. If you don't have a compose spec or prefer to follow our API, please see [here](https://api.omnistrate.cloud/docs/external/) ## Does my data get sent to your servers? Only what you choose to send. The data plane (aka your application) runs on your infrastructure or your customers' infrastructure. For example, when your customers interact with your application, none of the data is sent to us or runs through our infrastructure unless you deploy your application using Hosted SaaS. We only receive encrypted metadata needed to serve your control plane operations, and any data you choose to display on observability dashboards. ## Can non-coders use this product? Yes - as long as you can build compose YAML and know the product requirements, you can create your SaaS without any coding background. ## Do you support any on-prem environment? Currently, we support major cloud providers. However, our architecture is extendable to on-prem environments and continuously working with a few partners to extend our offering. Please [contact us](https://www.omnistrate.com/contact) for more details and we would love to discuss more ## What's your execution environment? We use K8s underneath the scenes but easily extendable to non-K8s environment. ## I have my own operator, can I integrate with Omnistrate? Please see [this page](https://docs.omnistrate.com/usecases/integrate-existing-stack/#custom-application-management-layer) ## I have my helm charts, can I integrate with Omnistrate? We are working on extending the helm support, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) for more details ## Do you have any standard compliance? Does it mean we will also be compatible? We are SOC2 type I/type II compliant. We are looking to extend it to other standards. If you have other needs, please reach out to us at [support@omnistrate.com](mailto:support@omnistrate.com) for more details Yes - by definition, your control plane will also be compliant. While we help accelerate your compliance for your SaaS, you will still need to establish compliance for data plane and corresponding infrastructure, internal processes, and team. ## How do you ensure reliability of my SaaS during major disaster? We have a decoupled design to avoid any cross-region dependencies for your SaaS control plane. In addition, we take constant backups to protect any necessary metadata. Please [contact us](https://www.omnistrate.com/contact) if you have any specific questions.