VCF 9.1 Enable High Availability for a Small VCF Deployment for VCF Management Services (VCFMS)

I still working on my shutdown en startup script for VCF 9.1

Because it is a small lab you have only 1 control plane and 3 workers.
A option to have second control plane vm would be nice for easier recovery.


Then I read about: William Lam: VCF 9.1 – Enabling High Availability for a Small VCF Management Services (

VCF

) Deployment

Tried it:

Checked operations.

Checked vCenter

I have now also logs running in VCF management services.
Go back to a single control plane is the same way disable HA.

When I have some time i want to test: Leaha’s Blog: VCF 9 Management Services Lab Downsize

VCF 9.1 Critical JV Certificate Caching Bug Can Cause Silent Outage

A field experience with VCF Operations and Fleet Lifecycle Management

The problem I encountered

Over the last few weeks, I encountered a strange issue in VMware Cloud Foundation 9.1. The Build > Software and Build > Lifecycle functionality was no longer working correctly.

Figure 1. Lifecycle functionality failing in the VCF interface.

My first thought was that something had gone wrong with the VCF Services Runtime control plane. I did not immediately know why it had failed, and unfortunately I did not have a VM-level backup of the VCF Services virtual machines available.

After investigating the symptoms, I suspected that I had run into a VCF 9.1 certificate-related bug.

Update: critical certificate caching issue

VCF 9.1 contains a critical issue in which the Java Virtual Machine (JVM) caches an internal security certificate during the initial bootstrap.

Issue details

Although the system correctly rotates the certificate in the background, the application does not dynamically reload the renewed certificate. Because these certificates have a strict 90-day validity period, environments deployed at General Availability (May 12, 2026) can begin experiencing silent outages around August 10, 2026.

Impact

The failure can occur without advance health-check warnings. Once the 90-day certificate threshold is reached, the environment may experience:

• Generic “503 Service Unavailable” errors
• Loss of access to affected UI functionality and blockage of component deployments
• Failure of day-two operations within Fleet Lifecycle Management (Fleet LCM)

What this looked like in my environment

In my case, the symptoms appeared in the VCF Build and Lifecycle areas and initially looked like a failure of the VCF Services Runtime control plane. The certificate-caching issue provides a plausible explanation for this behavior.

Action required

 

If you are running VCF 9.1, review the relevant Broadcom Knowledge Base articles and apply the documented workaround for the affected component. The source document identifies the following KB topics:

Takeaway

A certificate can be successfully rotated on the platform while an application continues using an older certificate cached by its JVM. That makes this issue particularly difficult to recognize: certificate rotation may appear healthy even though application communication eventually starts failing.

For VCF 9.1 environments approaching or exceeding 90 days since deployment, certificate-related TLS errors and unexplained 503 responses should therefore be investigated promptly.

VCF LOG Deployment changed in 9.1.0.400

Since VCF 9.1.0.400, a lot has changed in the platform and one of the more noticeable shifts is how log management works. It is no longer a standalone virtual appliance.

Instead, it now runs as a service inside the VCF Management Services Cluster, the same Kubernetes-based runtime that also hosts Fleet Management and the other core management components.

Additionally, in this blog I walk through how to deploy VCF Log Management in 9.1.0.400.

What Changed

In VCF 9.0 and earlier, Log Management (formerly Aria Operations for Logs) was deployed as its own appliance VM or appliance cluster. Additionally, from 9.1.0.400 onward, it is a built-in, containerized core service that runs natively on the VCF Management Services Platform. Deployment, lifecycle management, configuration, access control, and alerting are now all handled centrally through VCF Operations, rather than through a separate product UI.

Prerequisites

• DNS: A DNS-registered FQDN (with both forward and reverse lookup working)
• Capacity in the VCF Services Runtime: Deploying Log Management can trigger an automatic resize of the underlying Kubernetes cluster, so plan for the additional compute ahead of time.
• Sizing decision: The deployment size you choose determines how much compute and storage each node gets.
• Networking: Deployment currently supports VLAN-backed networks for the VCF Management Services; overlay networking is not yet supported for this component.

Resource Requirements for Log Management Deployments.

Size

vCPUs per Node

Memory per Node

Storage per Node

Number of Nodes

Small

8

16 GB

575 GB

1-19

Medium

24

48 GB

575 GB

3-19

Large

32

64 GB

575 GB

3-19

Procedure

1.Log in to the VCF Operations user interface at https://<vcf_operations_fqdn
2.Navigate to Build Lifecycle VCF Management.
3.On the Components tab, select Add Component Log management and configure the parameters for the deployment.


Now click Next to proceed to the Summary page for final review before deployment. Click Finish to start the deployment. As soon as the Deployment is done, you will see Log Management is running under VCF Operations -> Build -> Lifecycle.

Additionally, you can find the logs under VCF Operations -> Operate -> Log..

If you click on Log Sources, you will get an overview of the currently configured integrations, but this does not mean, that the Log collection itself is enabled.

Additionally, navigate to VCF Operations -> Operate -> Administration -> Configurations -> Log Collection to check Status. However, under Status, the Log collection is not enabled.

If you click on the “i” under Status you will see:

This is a little bit strange in my opinion, but yes, you have to switch back to the Integrations page inside VCF Operations Operate Administration Integrations and Edit your VMware Cloud Foundation Account:

From the Account page switch to Domains:

Next, on the Domain Page, start with vCenter to activate Logs Collection by clicking ‘Activate Log Collection’ and SAVE.

Same procedure i.e. NSX:

Additionally, after activiation of Log collection you will see more Logs inside VCF Operations:


.

Next, go to VCF Operations Operate Administration Configuration Log Collection.
Next, click Edit and scroll down.

Additionally, enable (Edit Log collection configuration).

.

What You Can Learn from a Minimum Resources 2 Node VCF 9 Lab Deployment to a Real Scenario

After three great sessions at VMUG Connect Amsterdam and VMUG Connect Online, and VMUGNL Summer Sessions. I’ve had requests to share the slides for my session “What You Can Learn from a Minimum Resources 2 Node VCF 9 Lab Deployment to a Real Scenario“.

The response was great.

In the slide deck you will find:

✅ The Start of My VCF Story
✅ First Challenge
✅ Why i wanted to build a real Senario
✅ Why i choise going to a build a real VCF scenario
✅ The Road
✅ The Lessions Learned & Next Steps

VCF 9.1: Fixing Root Account Password Expiration Issues

.

One of the first post-deployment tasks after installing VMware Cloud Foundation (VCF) 9.1 is configuring a password policy to ensure compliance with your organization’s security standards.



While the password policy is successfully applied to most managed accounts, you may notice that the root accounts of the VCF Operations appliance and the VCF Proxy appliance do not follow the configured password expiration policy.


As a result, the expiration date shown in VCF Management remains unchanged, even though a password policy has been configured.

Symptoms

You may observe one or more of the following:

• A password policy is configured successfully in VCF Management.

• Compliance checks complete without errors.

• The root account of the VCF Operations appliance still shows an incorrect or outdated password expiration date.

• The same behavior occurs on the VCF Proxy appliance.


This can be confusing because the password policy appears to be configured correctly, but it is not enforced for these Linux root accounts.

Why Does This Happen?

The password policy configured in VCF does not automatically update the Linux root account password aging settings on the VCF Operations and Proxy appliances.

Instead, these appliances continue to rely on the Linux ‘chage’ configuration to determine when the root password expires.


VMware has documented this behavior and provided a straightforward workaround.

https://knowledge.broadcom.com/external/article/441344/configured-password-policy-is-not-being.html

Resolution

Before making any changes, enable SSH access on the VCF Operations appliance if it is currently disabled.
https://knowledge.broadcom.com/external/article/315976/enabling-ssh-access-in-aria-operations.html

Step 1 – Connect to the VCF Operations Appliance

SSH to the VCF Operations appliance using an administrative account.

ssh admin@<vcf-operations-appliance>

Step 2 – Verify the Current Password Expiration

Run:

sudo chage -l root

Step 3 – Configure the Password Expiration

Configure a 365-day password lifetime:

sudo chage -M 365 root

Step 4 – Verify the Change

Run:

chage -l root

Step 5 – Repeat for the VCF Proxy Appliance

Repeat the same commands on the VCF Proxy appliance, either through SSH (if enabled) or via the VMware console.

Wait for VCF to Update

Notes

The updated password expiration date is not reflected immediately in the VCF Management interface.

VCF periodically refreshes password information, so it may take 10–30 minutes (or longer depending on your environment) before the new expiration date appears.

.

• This change only affects the Linux root account.
• Verify the setting after appliance upgrades.

• Adjust the password lifetime (90, 180, 365 days, etc.) according to your security policy.

Conclusion

Although VCF 9.1 allows administrators to centrally configure password policies, the Linux root accounts on the VCF Operations and VCF Proxy appliances continue to rely on the local Linux password aging configuration.

Updating the password expiration with the chage command ensures that the root account complies with your organization’s password policy. Once VCF completes its next inventory synchronization, the correct expiration date is displayed in the VCF Management interface.

VCF 9.0 Automate VMware Cloud Foundation Startup and Shutdown with PowerCLI

Powering your VMware Cloud Foundation “lab” environment on and off shouldn’t be a manual process.

A complete shutdown of a VMware Cloud Foundation (VCF) environment is uncommon, but for some energy savings and some time you does not use your lab often, you want a repeatable, reliable, and automated procedure, manually powering dozens of virtual machines in the correct order is both time-consuming and error-prone.

To solve this problem, I created two lightweight PowerCLI scripts:

• VCF 9.0 Small Startup Script
• VCF 9.0 Small Shutdown Script

.

vCenter and ESX hosts are manual (For vSAN cluster I did not find the correct code yet!)
Let me know if you have any questions or addons

Both scripts are available on GitHub and are designed to automate the startup and shutdown of a VMware Cloud Foundation management domain.

GitHub Repository

.

Note: Both Scripts do not work with VCF 9.1!!!!

.

Why These Scripts?

Although VMware Cloud Foundation automates the deployment and lifecycle of the platform, a full platform shutdown still requires the administrator to respect service dependencies.

For example:

• Domain Controllers must be available before authentication works.
• DNS must be online before many VMware services can resolve hostnames.
• vCenter must be operational before SDDC Manager can communicate with the infrastructure.
• NSX components depend on both networking and vCenter.
• VCF services should only start after the management platform is healthy.

Powering everything on simultaneously often results in services that need additional time—or even manual intervention—to recover.

These scripts automate the entire sequence descripted als following:

Start the Management Domain (VCF 9.0)’

Shut Down the Management Domain (VCF 9.0)

Typical Use Cases

These scripts are useful in many environments, including:

• Home labs
• Demonstration environments
• Disaster Recovery testing
• UPS maintenance
• Complete datacenter power outages
• Scheduled maintenance windows
• Hardware replacements

I personally use them in my VCF lab, where powering the environment up or down manually became repetitive and unnecessarily time-consuming. Automating the sequence not only saves time but also ensures a consistent and predictable startup every time.

Customizing the Scripts

Every VMware Cloud Foundation deployment is different.

The scripts are intentionally straightforward so you can easily adapt them by:

• Changing the startup order
• Adding custom virtual machines
• Removing components you don’t use
• Increasing wait times
• Adding health checks
• Integrating notifications
• Extending the logging

Because everything is written in PowerCLI, modifications are simple and require only basic scripting knowledge.

Future Improvements

Some ideas I’m considering for future releases include:

• Automatic dependency discovery
• Email notifications
• Automatic service validation
• Parallel startup where dependencies allow

Contributions and suggestions from the community are always welcome.

Lessons Learned

During development I discovered:

• VMware Tools are the best indicator that a guest OS is ready.
• Fixed sleep timers are unreliable because boot times vary.
• Starting all VMs simultaneously doesn’t necessarily reduce the total startup time.
• Graceful shutdowns significantly reduce recovery issues.
• Simplicity makes the scripts easier to customize.

Download

You can download the latest version from my public GitHub repository:

You can also find more VMware Cloud Foundation automation projects, PowerCLI scripts, and lab guides on my website:

Conclusion

A VMware Cloud Foundation environment consists of many interconnected services, and those services should be started and stopped in the correct order.

These lightweight PowerCLI scripts automate that process, making startup and shutdown predictable, repeatable, and significantly less error-prone. Whether you’re running a production management domain or a small VCF lab, automating these operational tasks saves time and reduces the risk of mistakes.

If you have ideas for improvements or additional features, feel free to open an issue or submit a pull request on GitHub. Happy automating!

.

Config a VCF (vSAN ESA) host the Easy Way

A while ago i created: 

V1: Config vSAN ESA host or VCF ESA vSAN Host the easy way with Config-VSAN-ESA-VCF-Lab-Host Script.

V2: Config a VCF (vSAN ESA) host the Easy Way

With the release of VCF 9.1 is now time for again a updated version!

What does the script now:

✅ Disable ipv6

✅ Set DNS domain name

✅ Rename local datastore

✅Configure NTP

✅ MTU 9000

✅ Installs the vSAN ESA Hardware Mock VIB &

✅ Workaround to reduce impact of resync traffic in vSAN ESA clusters utilizing a 10G network

✅ Installs the Synology NFS Plug-in for VMware VAAI

✅ Memory Optimalisation Additional Transparent Page Sharing management capabilities and new default settings

 ✅ AMD Zen4/Zen5 IPMI Thermal Driver for ESX AMD Zen4/Zen5 IPMI Thermal Driver for ESX Fling

✅ Generate new certificate on the ESXi host (for the VCF verification check)

✅ Ask are you running Miniforum MS-A2(AMD) host & Optimalization see
VCF 9.1 – Comprehensive ESX Configuration Workarounds for Lab Deployments (Except the vSAN Compression Algorithm)

 ✅ Enable Memory Tiering for 9.1 and 9.0 and older (Filter OS Disk)

vSAN ESA Mock and the AMD Zen4/Zen5 IPMI Thermal Driver can be download on the Broadcom Fling page.

(https://support.broadcom.com/group/ecx/free-downloads) section and select Flings (https://support.broadcom.com/group/ecx/productdownloads?subfamily=Flings&freeDownloads=true)

You need to download the vibs separately!
For the installs put the vib’s in the same map as the script
.

You can download the script: HERE

.

VCF 9.1 What fixed the stuck deployment at “Deploy and Configure VCF Management Platform”

Overview

While deploying VMware Cloud Foundation (VCF) 9.1 in a homelab environment, the installation repeatedly failed during the ‘Deploy and Configure VCF Management Platform’ stage. Despite performing nine completely clean installations, the deployment consistently stopped at the same point.

Error Observed

The deployment task failed with the message: ‘Add VM Name Prefix to NSX Firewall Exclusion List’. The failure was identified in /var/log/vmware/vcf/domainmanager/domainmanager.log

Initial Research

Several Broadcom Knowledge Base articles appeared relevant, including KB440449 and KB 441122. Although the symptoms were similar, neither article fully resolved the issue.

VMSP Configuration Review

The original VMSP configuration used a name value matching the prefix of the fleetFqdn. The configuration was modified to use a unique VMSP cluster name. While this appeared promising, the issue persisted.

Additional Troubleshooting

Additional troubleshooting included changing VMSP IP ranges, rebuilding DNS records, validating forward and reverse DNS resolution, and reviewing deployment logs for networking issues.

Root Cause Analysis

The issue was ultimately not caused by the NSX firewall exclusion configuration. Multiple infrastructure issues contributed to deployment instability.

Resolution

1. Configure a single authoritative NTP source, preferably the domain controller.
See the planning and preparation workbook
https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/planning-and-preparation.html
2. Verify DNS records and name resolution.
3. Upgraded to a dedicated 10G Switch Ubiquiti UniFi Pro XG 8 PoE ipv 2.5GB Ubiquity Switch
Switch Pro XG 8 PoE - Ubiquiti Store Europe
4. Replace faulty network components.

Conclusion

Although the deployment failure appeared to indicate an NSX firewall exclusion issue, the underlying cause was network instability combined with infrastructure configuration problems. After correcting NTP configuration, validating DNS, upgrading network connectivity, and replacing the defective SFP+ module, the VCF 9.1 deployment completed successfully.

What’s New in Private Cloud: VMware VCF 9.1 Enhancements

VMware Cloud Foundation (VCF) 9.1 is here — and it’s one of the most feature‑packed releases in years. This update isn’t just incremental; it’s a strategic modernization of compute, storage, networking, security, and operations across the entire private cloud stack.


.

Let’s break down the biggest enhancements and why they I think they matter.

Modernizing Infrastructure Economics with vSphere Foundation 9.1

VCF 9.1 brings several powerful updates to the vSphere layer, aimed at improving performance efficiency and reducing operational overhead.

Enhanced NVMe Memory Tiering

Workloads that demand high throughput and low latency benefit from smarter memory tiering. NVMe-based memory tiers now deliver improved performance and flexibility. (And yes — many are hoping Secure Boot support lands here as well.)


Parallel Processing of DRS vMotion

DRS can now process multiple vMotions in parallel, dramatically reducing cluster balancing times. This is especially impactful in large-scale environments with frequent workload mobility.

Live Patching for TPM-Enabled Hosts

Live patching now works even on hosts with TPM enabled — a huge win for security-conscious organizations that previously had to choose between uptime and compliance.

Networking Updates: Scale, Simplicity, and Smarter Automation

VCF 9.1 introduces major networking enhancements that streamline operations and expand connectivity options.

Enhanced Day-2 VM Lifecycle Management

Networking changes for VMs — including NIC updates, IP changes, and security policies — are now easier and more automated.

Existing VLAN Connectivity via Distributed Transit Gateways

You can now bridge existing VLAN-based networks into VCF environments more seamlessly, reducing migration friction and simplifying hybrid designs.

Streamlined Firewalls & Automated Inter-VPC Security

Security policies between VPCs are now automated, reducing manual rule creation and improving consistency across tenants.

Terraform Provider Enhancements

Better support for tenant-level policy and content management means more automation and cleaner IaC workflows.

Simplified Workload Connectivity & Enhanced Network Scale

EVPN-VXLAN Interoperability

VCF 9.1 now supports EVPN-VXLAN interoperability with the physicalnetwork fabric. This is a major step toward fully integrated, fabric-aware cloud networking.

Network Assessment & VPC Planning

New tools and workflows help architects plan VPC layouts, assess network readiness, and avoid misconfigurations before deployment.

Optimize, Modernize & Protect Storage with vSAN in VCF 9.1

Storage gets a significant upgrade in this release, especially for environments focused on efficiency and resilience.

Encryption for vSAN Global Deduplication

Global dedupe is now compatible with data-at-rest encryption — a long-awaited capability for secure, space-efficient storage.

Enhanced Stretched Cluster Capabilities

Improved resilience and smarter failure handling strengthen business continuity for mission-critical workloads.

Automated Storage Policy Management

Policies now adjust automatically based on cluster configuration changes, reducing manual tuning and risk of misalignment.

Strengthening Zero Trust Security & Platform Resilience

Security is a major theme in VCF 9.1, with improvements across the stack.

Data-at-Rest Encryption for Global Dedupe

This ensures encrypted storage without sacrificing dedupe efficiency — a rare combination in enterprise storage.

Quick Patching for vCenter

Faster patch cycles reduce exposure windows and simplify maintenance.

Live Patching for TPM-Enabled Hosts

As mentioned earlier, this is a major operational win for secure environments.

Continuous Compliance & Integrated Cyber Recovery

VCF 9.1 pushes deeper into automated compliance and recovery workflows.

Compliance Monitoring & Desired State Remediation

The platform now continuously checks VCF components against desired state and can automatically remediate drift.

VPC Policy-Based Connectivity

Security and connectivity policies can now be applied consistently across VPCs, improving governance and reducing misconfigurations.

VMware Data Services Manager 9.1: Modern Databases for AI & Cloud

Microsoft SQL Server 2022 Now GA

SQL Server 2022 is now fully supported and generally available through DSM 9.1, enabling automated lifecycle management for modern database workloads — including those powering AI and analytics.

Want to See It in Action?

VMware has published a full VCF 9.1 video podcast series that dives deeper into the new capabilities:

Enough to do in my Homelab Starting with Upgrade and testing the new features!!

Running VCF 9 in a Nested Lab using vSAN ESA on a single host is that not great?

For the deployment I used a server that has 16 cores en 32 cores total and a lot of ram!

Installation on Nested ESX VM’s

For the nested ESX VMs I am using the generic ESX image. In my lab I’m using 4 nested ESX nodes for vSAN ESA (Minimum is needed is 3).

  • Ideally add 24 vCPUs to the nested ESX, but VCF Automation is capable of being deployed with 16 vCPUs
  • Check Expose hardware assisted virtualization to the guest OS under CPU
  • Add a NVME controller and connect disks to this controller if planning to use vSAN ESA (Don’t forget to remove the old Controller)
  • Add at least two VMXNET3 nics connected to the same network
  • Choose VMware ESXi 8.0 or later as the Guest OS version.
  • Change the IDE CD-ROM to SATA ROM (Failed to locate kickstart on Nested ESXi VM CD-ROM in VCF 9.0)

Afbeelding met tekst, schermopname, nummer, LettertypeDoor AI gegenereerde inhoud is mogelijk onjuist.

Config the 4 Host for vSAN ESA

To config the 4 Hosts for vSAN ESA I used my own powercli script where I blogged about here: Config vSAN ESA host or VCF ESA vSAN Host the easy way with Config-VSAN-ESA-VCF-Lab-Host Script

You can download the script HERE!

Installing the Vib works still great. No problem with ESA  Pre-check: ✅

If you have problems check this great blog from William Lam: vSAN ESA Disk & HCL Workaround for VCF 9.0

VCF installer.

At first you have to deploy the VMware Cloud Foundation

The name of the installer is a little off: VCF-SDDC-Manager-Appliance-9.0.0.0.24703748.ova

Download token

If you don’t have a download token from the Broadcom support then it’s a little complicated. I won’t go into this in depth now but here is a nice article if you have a Synology Nas: VCF 9.0 Offline Depot using Synology

(Note: For any vExpert reading this… You need to claim your own domain name because build your profile is not working with a @hotmail.com or @gmail accounts, Noted by Broadcom)

Afbeelding met tekst, schermopname, Parallel, ontwerpDoor AI gegenereerde inhoud is mogelijk onjuist.

I looks something like this!
Afbeelding met tekst, schermopname, LettertypeDoor AI gegenereerde inhoud is mogelijk onjuist.

10GBE Nic Pre-Check Issue

Next thing to do was: Disable 10GbE NIC Pre-Check in the VCF 9.0 Installer or this should also work: How to change the vmxnet3 link speed of a VM (Not tested)

Afbeelding met tekst, elektronica, schermopname, softwareDoor AI gegenereerde inhoud is mogelijk onjuist.

Deployment

I have deployed the beta in a earlier state. For deployment I used the same json file to start with.

VCF Automation will deployed automatically that was only the change that I changed it

Afbeelding met tekst, schermopname, nummer, softwareDoor AI gegenereerde inhoud is mogelijk onjuist.

Afbeelding met tekst, schermopname, Website, ontwerpDoor AI gegenereerde inhoud is mogelijk onjuist.

It’s running!
Afbeelding met tekst, schermopname, scherm, nummerDoor AI gegenereerde inhoud is mogelijk onjuist.

Afbeelding met tekst, lijn, nummer, LettertypeDoor AI gegenereerde inhoud is mogelijk onjuist.

Performance running Nested!!

Afbeelding met tekst, lijn, Perceel, LettertypeDoor AI gegenereerde inhoud is mogelijk onjuist.

Above the screenshot from the host with de Nested VCF lab deployment it’s quite cpu intensive (2 x CPU 8cores / 16 Threads)

Optimalization

After the deployment I did three things
– On the vSAN ESA Cluster (Disable Auto Disk Claim) (Keep the warning away)
– On the NSX VM reduce the cpu reservation from 6000 🡪 2000 (It helps but not enough)
– VCF automation VM 24 cores to 16 cores.

NSX en VCF automation are really CPU intensive.

I had deployed also two Edge servers but the server did not like that. Edges are also CPU intensive.

I am thinking about adding MS-01 or MS-A2 for splitting some load.

Translate »