banner

Marcelo Gigliotti

Principal Product Manager

2026-10-06T00:00:00.000Z
eds-infoscale:,eds-infoscale:tags/infoscale

AI Can Understand Your Kubernetes Environment. But Can It Keep Your Applications Running?

OpenShift Lightspeed brings AI-assisted intelligence to Kubernetes operations. InfoScale adds the data availability, application mobility and state protection needed to turn that intelligence into operational resilience.

AI is changing how enterprises operate infrastructure. With tools such as Red Hat OpenShift Lightspeed, administrators can increasingly interact with complex Kubernetes environments in natural language. Instead of moving between dashboards, documentation, command-line tools and runbooks, teams can ask questions directly:

Which worker nodes have capacity?

What storage configuration is best for this workload?

Is this application ready to move?

Where is there risk in the environment?

Getting an intelligent answer in seconds instead of spending 30 minutes piecing one together is valuable. But for enterprises running mission-critical applications, answering the question is only the beginning.

The bigger challenge is what happens next. If a node fails, can the application move without losing access to its data? If infrastructure conditions change, can protection adapt with the workload? If an application restarts somewhere else, is its data still valid and consistent? If a virtual machine moves across infrastructure, can its data and operational context move with it? And when disruption happens, can the environment preserve what the application needs to keep the business running? These aren't simply Kubernetes or storage questions.

They're questions of operational resilience.

And answering them requires protecting more than infrastructure.

It requires protecting the data and operational state of the application.

Kubernetes manages desired state. The business depends on operational state.

Kubernetes transformed infrastructure by making it dynamic. A container disappears? Restart it. A node fails? Reschedule the workload. Demand increases? Scale the application.

That model works because Kubernetes continuously compares the environment against a defined desired state and works to reconcile the two. But a running enterprise application depends on another kind of state. Its operational state.

Operational state is the live combination of application context, persistent data, infrastructure, configuration, dependencies and relationships required for an application to operate correctly. For a stateless service, replacing compute may be enough. For a database, payment platform, ERP system, AI workload or other stateful application, it isn't. The application needs its data. That data needs to be available and consistent.

The infrastructure running the workload needs access to it. Dependencies and relationships need to remain intact. And all of this needs to survive as infrastructure changes underneath the application. That's the distinction at the center of modern operational resilience:

Kubernetes maintains the desired state of the platform.

Operational resilience protects the data and operational state required to keep the application running.

Why data changes the resilience equation

Enterprise infrastructure is becoming increasingly disposable. Nodes can be replaced. Containers can restart. Virtual machines can migrate. Clusters can be rebuilt. Workloads can move between environments. Data is different. A new Kubernetes worker can be created in minutes. You can't recreate yesterday's transactions by provisioning another node. That means the architecture underneath stateful applications has an outsized impact on resilience.

Consider a workload experiencing a sudden resource spike. The infrastructure layer may determine that another worker node has significantly more CPU and memory available. Moving the workload seems straightforward.

But there is another question:

Can the application's data move with it?

In architectures where application mobility depends on the physical location of a storage replica, the best destination from an infrastructure perspective may not actually be available to the application. The compute decision and the data decision become two different problems.

InfoScale takes a different approach.

Its software-defined architecture does not tie application mobility to the location of an individual storage replica. That means the worker node with the right resources can also be a viable destination for the workload without requiring the application to remain pinned to wherever a particular data copy resides.

Infrastructure can change. The application and its data remain connected. That is a fundamental requirement for operational resilience.

AI gives administrators a new way to understand operational state

This is where OpenShift Lightspeed makes the story particularly interesting. OpenShift Lightspeed brings an AI assistant directly into the OpenShift console, allowing administrators to ask questions about their environment using natural language. InfoScale for Kubernetes already exposes its operational information through Kubernetes-native objects, including StorageClasses, PVCs and InfoScale CRDs.

That means there doesn't need to be an entirely separate layer just to make InfoScale visible to AI. The objects doing the work are also capable of providing the context. Lightspeed can query Kubernetes state and help an administrator reason about what is happening across the environment. InfoScale provides the data and resilience architecture underneath the application. Together, they demonstrate something bigger than an AI assistant for storage administration.

They show what happens when AI-assisted understanding meets state-aware infrastructure. And that creates the foundation for a new model of resilience.

Understand. Protect. Adapt. Recover.

Traditional disaster recovery typically starts after something has already gone wrong. A failure occurs. Someone detects it. The team investigates. A recovery decision is made. A runbook is executed. Infrastructure is restored. Data is recovered. The application is validated. Eventually, the business is operational again. Every step takes time.

And every manual decision creates another opportunity for delay. Autonomous operational resilience changes the model.

Instead of treating resilience as a recovery process that begins after disruption, the environment can continuously understand what is happening and maintain the data and operational state required to respond.

That model can be thought of in four steps:

Understand the operational state.

Protect the data and state that matter.

Adapt as conditions change.

Recover the application when disruption can't be absorbed.

OpenShift Lightspeed and InfoScale demonstrate what those principles can look like in practice.

Understand: Turn infrastructure complexity into context

It's 2 a.m.

A new data disk needs to be provisioned before the batch job that closes yesterday's transactions begins.

The administrator knows the workload will be write-heavy but doesn't remember the exact storage layout previously used for a similar database.

The traditional workflow is familiar.

Find the runbook.

Search documentation.

Look for the previous StorageClass.

Ask someone in Slack.

Or make the best decision you can from memory.

With OpenShift Lightspeed, the administrator can instead ask:

Assume an InfoScale CSI StorageClass with these parameters: layout, fstype, ncol and stripeunit. For a write-heavy PostgreSQL data disk with high sequential I/O, what values would you recommend for layout, ncol and stripeunit?

Lightspeed can reason over storage engineering practices and return the recommendation using the same parameters InfoScale consumes.

The administrator remains in control of the action.

But the time between question and informed decision can shrink dramatically.

This is the first piece of autonomous operational resilience:

understanding the environment quickly enough to act before a condition becomes a disruption.

Protect: Keep the data and operational state valid

Intelligence alone doesn't create resilience.

The underlying application state still has to be protected.

Consider a worker node that suddenly becomes unreachable while hosting a stateful application.

Kubernetes can identify the failure and restart the workload elsewhere.

But before another instance begins accessing the application's persistent storage, the environment needs confidence that the failed node can no longer write to that data.

Otherwise, two systems could attempt to modify the same storage.

That creates the possibility of conflicting writes and data corruption.

InfoScale uses I/O fencing to protect persistent data during these scenarios, preventing an inaccessible node from continuing to access protected storage so that the workload can safely resume elsewhere.

That's more than storage availability.

The infrastructure changed.

The application moved.

But the integrity of the application's operational state was preserved.

The goal isn't simply to restart the workload.

The goal is to restart it in a state the business can trust.

Adapt: Let applications move as conditions change

Not every disruption begins with a failure.

Sometimes the environment simply changes.

A worker node reaches 90% CPU during a traffic spike.

Another node has significantly more capacity.

The workload needs to move before resource pressure becomes customer-facing downtime.

An administrator can ask OpenShift Lightspeed:

Which worker nodes have the lowest current CPU and memory utilization, and are they eligible to receive a live migration from the current worker node?

Lightspeed can correlate the available infrastructure information and help identify an appropriate destination.

But identifying the best node doesn't necessarily mean the application can move there.

Storage architecture matters.

With replica-dependent architectures, migration may be constrained by where a valid copy of the workload's data currently resides. The best compute destination may not be a valid storage destination.

InfoScale removes that dependency.

Because application mobility isn't tied to the location of an individual storage replica, infrastructure decisions can be based on the needs of the workload rather than the physical location of a particular data copy.

That changes the role of storage.

Instead of constraining where an application can run, the data layer helps enable the application to adapt as infrastructure conditions change.

That's an important step toward autonomous resilience:

Protection should follow the application—not restrict it.

Optimize: Change the environment without disrupting the business

Operational state also changes over time.

A PVC provisioned eight months ago for a prototype may now support a production order-processing application.

The original storage configuration may no longer match the workload's actual I/O pattern.

Performance gradually deteriorates.

Everyone knows the configuration should probably change.

But changing production infrastructure traditionally creates another problem: downtime.

So the imperfect configuration survives because disrupting the business feels riskier than tolerating the performance problem.

AI-assisted operations can shorten the first part of that process.

An administrator can ask Lightspeed what storage layout is better suited to the workload and receive recommendations mapped directly to InfoScale's StorageClass parameters.

InfoScale then provides the underlying ability to evolve the storage supporting the application.

The broader principle is important:

Operational resilience shouldn't require applications to remain frozen in yesterday's architecture.

Applications change.

Usage changes.

Infrastructure changes.

Protection needs change.

A resilient environment needs to adapt with them.

Recover: Preserve state across clusters and sites

Eventually, some disruptions cannot be absorbed locally.

A cluster can become unavailable.

A site can fail.

A cyber incident can force workloads into another environment.

At that point, the problem is larger than restarting a container.

The application needs to be reconstructed somewhere else with the data and resources required to operate.

InfoScale for Kubernetes provides capabilities for replication and disaster recovery across environments through components including DataRep, DRPlan and GCM.

For OpenShift Virtualization workloads, that can mean preserving and restoring the persistent data and resources required to bring a virtual machine back into operation in another environment.

Again, the objective isn't merely:

Did we recover the data?

The more important question is:

Did we recover the application into an operational state?

That distinction is increasingly important as enterprises move mission-critical virtual machines and stateful applications onto Kubernetes platforms.

From AI-assisted operations to autonomous operational resilience

Today, OpenShift Lightspeed provides an AI-assisted interface for understanding and reasoning about the environment.

The administrator asks.

The system answers.

The administrator decides what happens next.

That's valuable on its own.

But it also points toward something larger.

Imagine an environment that continuously understands:

The objective isn't to hand every infrastructure decision to an AI model.

It's to combine intelligence, policy, automation and state-aware infrastructure so resilience can happen with dramatically less manual intervention.

That is the progression from AI-assisted operations to autonomous operational resilience.

What autonomous operational resilience looks like

The traditional resilience model is largely reactive:

Disruption → Detect → Diagnose → Decide → Recover → Validate

The emerging model is continuous:

Understand → Protect → Adapt → Validate

And when recovery becomes necessary:

Recover the application and its operational state.

That difference matters because recovery time isn't only determined by how quickly infrastructure can restart.

It's determined by how many decisions humans need to make before the business can operate again.

If AI can accelerate understanding, automation can accelerate action, and InfoScale can preserve the data and operational state required by the application, organizations can reduce the gap between infrastructure disruption and business continuity.

The enterprise resilience conversation needs to move beyond infrastructure

For CIOs, infrastructure leaders, platform engineering teams and enterprise architects, this changes the questions worth asking.

Don't ask only:

Is our Kubernetes platform highly available?

Ask:

Can our critical applications remain operational when the infrastructure underneath them changes?

Don't ask only:

Do we have multiple copies of our data?

Ask:

Can the application access valid, consistent data wherever it needs to run?

Don't ask only:

Can Kubernetes restart the workload?

Ask:

Can it restart the workload in a state the business can trust?

Don't ask only:

Do we have a disaster recovery plan?

Ask:

Can we preserve and restore the data, application resources and relationships required to resume operations?

And don't ask only:

How can AI help us operate Kubernetes?

Ask:

What becomes possible when AI understands the environment and the infrastructure underneath it can protect and act on that state?

Protect the operational state

Infrastructure will fail.

Cloud services will change.

Nodes will disappear.

Applications will move.

Demand will spike.

Cyber incidents will happen.

Trying to eliminate every possible disruption isn't resilience.

Building systems capable of continuing through those disruptions is.

That's why the next phase of enterprise resilience isn't simply about making infrastructure more redundant or creating more copies of data.

It's about continuously understanding and protecting what the business actually depends on:

the application's data and operational state.

OpenShift Lightspeed demonstrates how AI can make complex infrastructure easier to understand.

InfoScale provides the data availability, application mobility, protection and recovery capabilities required to keep stateful applications operational as that infrastructure changes.

Together, they point toward a future where resilience becomes less dependent on an administrator finding the right dashboard, remembering the right command or executing the right runbook at exactly the right moment.

A future where systems can understand what is happening, protect what matters and adapt as conditions change.

That is autonomous operational resilience.

And for stateful enterprise applications, it starts by protecting the data and the state.