A field experience with VCF Operations and Fleet Lifecycle Management
The problem I encountered
Over the last few weeks, I encountered a strange issue in VMware Cloud Foundation 9.1. The Build > Software and Build > Lifecycle functionality was no longer working correctly.
Figure 1. Lifecycle functionality failing in the VCF interface.
My first thought was that something had gone wrong with the VCF Services Runtime control plane. I did not immediately know why it had failed, and unfortunately I did not have a VM-level backup of the VCF Services virtual machines available.
After investigating the symptoms, I suspected that I had run into a VCF 9.1 certificate-related bug.
Update: critical certificate caching issue
VCF 9.1 contains a critical issue in which the Java Virtual Machine (JVM) caches an internal security certificate during the initial bootstrap.
Issue details
Although the system correctly rotates the certificate in the background, the application does not dynamically reload the renewed certificate. Because these certificates have a strict 90-day validity period, environments deployed at General Availability (May 12, 2026) can begin experiencing silent outages around August 10, 2026.
Impact
The failure can occur without advance health-check warnings. Once the 90-day certificate threshold is reached, the environment may experience:
•Generic “503 Service Unavailable” errors
•Loss of access to affected UI functionality and blockage of component deployments
•Failure of day-two operations within Fleet Lifecycle Management (Fleet LCM)
What this looked like in my environment
In my case, the symptoms appeared in the VCF Build and Lifecycle areas and initially looked like a failure of the VCF Services Runtime control plane. The certificate-caching issue provides a plausible explanation for this behavior.
Action required
If you are running VCF 9.1, review the relevant Broadcom Knowledge Base articles and apply the documented workaround for the affected component. The source document identifies the following KB topics:
A certificate can be successfully rotated on the platform while an application continues using an older certificate cached by its JVM. That makes this issue particularly difficult to recognize: certificate rotation may appear healthy even though application communication eventually starts failing.
For VCF 9.1 environments approaching or exceeding 90 days since deployment, certificate-related TLS errors and unexplained 503 responses should therefore be investigated promptly.
Since VCF 9.1.0.400, a lot has changed in the platform and one of the more noticeable shifts is how log management works. It is no longer a standalone virtual appliance.
Instead, it now runs as a service inside the VCF Management Services Cluster, the same Kubernetes-based runtime that also hosts Fleet Management and the other core management components.
Additionally, in this blog I walk through how to deploy VCF Log Management in 9.1.0.400.
What Changed
In VCF 9.0 and earlier, Log Management (formerly Aria Operations for Logs) was deployed as its own appliance VM or appliance cluster. Additionally, from 9.1.0.400 onward, it is a built-in, containerized core service that runs natively on the VCF Management Services Platform. Deployment, lifecycle management, configuration, access control, and alerting are now all handled centrally through VCF Operations, rather than through a separate product UI.
Prerequisites
•DNS: A DNS-registered FQDN (with both forward and reverse lookup working)
•Capacity in the VCF Services Runtime: Deploying Log Management can trigger an automatic resize of the underlying Kubernetes cluster, so plan for the additional compute ahead of time.
•Sizing decision: The deployment size you choose determines how much compute and storage each node gets.
•Networking: Deployment currently supports VLAN-backed networks for the VCF Management Services; overlay networking is not yet supported for this component.
Resource Requirements for Log Management Deployments.
Size
vCPUs per Node
Memory per Node
Storage per Node
Number of Nodes
Small
8
16 GB
575 GB
1-19
Medium
24
48 GB
575 GB
3-19
Large
32
64 GB
575 GB
3-19
Procedure
1.Log in to the VCF Operations user interface at https://<vcf_operations_fqdn
2.Navigate to BuildLifecycleVCF Management.
3.On the Components tab, select Add Component Log management and configure the parameters for the deployment.
Now click Next to proceed to the Summary page for final review before deployment. Click Finish to start the deployment. As soon as the Deployment is done, you will see Log Management is running under VCF Operations -> Build -> Lifecycle.
Additionally, you can find the logs under VCF Operations -> Operate -> Log..
If you click on Log Sources, you will get an overview of the currently configured integrations, but this does not mean, that the Log collection itself is enabled.
Additionally, navigate to VCF Operations -> Operate -> Administration -> Configurations -> Log Collection to check Status. However, under Status, the Log collection is not enabled.
If you click on the “i” under Status you will see:
This is a little bit strange in my opinion, but yes, you have to switch back to the Integrations page inside VCF Operations Operate Administration Integrations and Edit your VMware Cloud Foundation Account:
From the Account page switch to Domains:
Next, on the Domain Page, start with vCenter to activate Logs Collection by clicking ‘Activate Log Collection’ and SAVE.
Same procedure i.e. NSX:
Additionally, after activiation of Log collection you will see more Logs inside VCF Operations:
.
Next, go to VCF Operations Operate Administration Configuration Log Collection. Next, click Edit and scroll down.
After three great sessions at VMUG Connect Amsterdam and VMUG Connect Online, and VMUGNL Summer Sessions. I’ve had requests to share the slides for my session “What You Can Learn from a Minimum Resources 2 Node VCF 9 Lab Deployment to a Real Scenario“.
The response was great.
In the slide deck you will find:
✅ The Start of My VCF Story ✅ First Challenge ✅ Why i wanted to build a real Senario ✅ Why i choise going to a build a real VCF scenario ✅ The Road ✅ The Lessions Learned & Next Steps
One of the first post-deployment tasks after installing VMware Cloud Foundation (VCF) 9.1 is configuring a password policy to ensure compliance with your organization’s security standards.
While the password policy is successfully applied to most managed accounts, you may notice that the root accounts of the VCF Operations appliance and the VCF Proxy appliance do not follow the configured password expiration policy. As a result, the expiration date shown in VCF Management remains unchanged, even though a password policy has been configured.
Symptoms
You may observe one or more of the following: • A password policy is configured successfully in VCF Management. • Compliance checks complete without errors. • The root account of the VCF Operations appliance still shows an incorrect or outdated password expiration date. • The same behavior occurs on the VCF Proxy appliance. This can be confusing because the password policy appears to be configured correctly, but it is not enforced for these Linux root accounts.
Why Does This Happen?
The password policy configured in VCF does not automatically update the Linux root account password aging settings on the VCF Operations and Proxy appliances. Instead, these appliances continue to rely on the Linux ‘chage’ configuration to determine when the root password expires. VMware has documented this behavior and provided a straightforward workaround. https://knowledge.broadcom.com/external/article/441344/configured-password-policy-is-not-being.html
SSH to the VCF Operations appliance using an administrative account. ssh admin@<vcf-operations-appliance>
Step 2 – Verify the Current Password Expiration
Run: sudo chage -l root
Step 3 – Configure the Password Expiration
Configure a 365-day password lifetime: sudo chage -M 365 root
Step 4 – Verify the Change
Run: chage -l root
Step 5 – Repeat for the VCF Proxy Appliance
Repeat the same commands on the VCF Proxy appliance, either through SSH (if enabled) or via the VMware console.
Wait for VCF to Update
Notes
The updated password expiration date is not reflected immediately in the VCF Management interface. VCF periodically refreshes password information, so it may take 10–30 minutes (or longer depending on your environment) before the new expiration date appears.
.
• This change only affects the Linux root account. • Verify the setting after appliance upgrades. • Adjust the password lifetime (90, 180, 365 days, etc.) according to your security policy.
Conclusion
Although VCF 9.1 allows administrators to centrally configure password policies, the Linux root accounts on the VCF Operations and VCF Proxy appliances continue to rely on the local Linux password aging configuration. Updating the password expiration with the chage command ensures that the root account complies with your organization’s password policy. Once VCF completes its next inventory synchronization, the correct expiration date is displayed in the VCF Management interface.
Powering your VMware Cloud Foundation “lab” environment on and off shouldn’t be a manual process.
A complete shutdown of a VMware Cloud Foundation (VCF) environment is uncommon, but for some energy savings and some time you does not use your lab often, you want a repeatable, reliable, and automated procedure, manually powering dozens of virtual machines in the correct order is both time-consuming and error-prone.
To solve this problem, I created two lightweight PowerCLI scripts:
•VCF 9.0 Small Startup Script
•VCF 9.0 Small Shutdown Script
.
vCenter and ESX hosts are manual (For vSAN cluster I did not find the correct code yet!) Let me know if you have any questions or addons
Both scripts are available on GitHub and are designed to automate the startup and shutdown of a VMware Cloud Foundation management domain.
Although VMware Cloud Foundation automates the deployment and lifecycle of the platform, a full platform shutdown still requires the administrator to respect service dependencies.
For example:
•Domain Controllers must be available before authentication works.
•DNS must be online before many VMware services can resolve hostnames.
•vCenter must be operational before SDDC Manager can communicate with the infrastructure.
•NSX components depend on both networking and vCenter.
•VCF services should only start after the management platform is healthy.
Powering everything on simultaneously often results in services that need additional time—or even manual intervention—to recover.
These scripts automate the entire sequence descripted als following:
These scripts are useful in many environments, including:
•Home labs
•Demonstration environments
•Disaster Recovery testing
•UPS maintenance
•Complete datacenter power outages
•Scheduled maintenance windows
•Hardware replacements
I personally use them in my VCF lab, where powering the environment up or down manually became repetitive and unnecessarily time-consuming. Automating the sequence not only saves time but also ensures a consistent and predictable startup every time.
Customizing the Scripts
Every VMware Cloud Foundation deployment is different.
The scripts are intentionally straightforward so you can easily adapt them by:
•Changing the startup order
•Adding custom virtual machines
•Removing components you don’t use
•Increasing wait times
•Adding health checks
•Integrating notifications
•Extending the logging
Because everything is written in PowerCLI, modifications are simple and require only basic scripting knowledge.
Future Improvements
Some ideas I’m considering for future releases include:
•Automatic dependency discovery
•Email notifications
•Automatic service validation
•Parallel startup where dependencies allow
Contributions and suggestions from the community are always welcome.
Lessons Learned
During development I discovered:
•VMware Tools are the best indicator that a guest OS is ready.
•Fixed sleep timers are unreliable because boot times vary.
•Starting all VMs simultaneously doesn’t necessarily reduce the total startup time.
A VMware Cloud Foundation environment consists of many interconnected services, and those services should be started and stopped in the correct order.
These lightweight PowerCLI scripts automate that process, making startup and shutdown predictable, repeatable, and significantly less error-prone. Whether you’re running a production management domain or a small VCF lab, automating these operational tasks saves time and reduces the risk of mistakes.
If you have ideas for improvements or additional features, feel free to open an issue or submit a pull request on GitHub. Happy automating!
While deploying VMware Cloud Foundation (VCF) 9.1 in a homelab environment, the installation repeatedly failed during the ‘Deploy and Configure VCF Management Platform’ stage. Despite performing nine completely clean installations, the deployment consistently stopped at the same point.
Error Observed
The deployment task failed with the message: ‘Add VM Name Prefix to NSX Firewall Exclusion List’. The failure was identified in /var/log/vmware/vcf/domainmanager/domainmanager.log
Initial Research
Several Broadcom Knowledge Base articles appeared relevant, including KB440449 and KB 441122. Although the symptoms were similar, neither article fully resolved the issue.
VMSP Configuration Review
The original VMSP configuration used a name value matching the prefix of the fleetFqdn. The configuration was modified to use a unique VMSP cluster name. While this appeared promising, the issue persisted.
Additional Troubleshooting
Additional troubleshooting included changing VMSP IP ranges, rebuilding DNS records, validating forward and reverse DNS resolution, and reviewing deployment logs for networking issues.
Root Cause Analysis
The issue was ultimately not caused by the NSX firewall exclusion configuration. Multiple infrastructure issues contributed to deployment instability.
Although the deployment failure appeared to indicate an NSX firewall exclusion issue, the underlying cause was network instability combined with infrastructure configuration problems. After correcting NTP configuration, validating DNS, upgrading network connectivity, and replacing the defective SFP+ module, the VCF 9.1 deployment completed successfully.
VMware Cloud Foundation (VCF) 9.1 is here — and it’s one of the most feature‑packed releases in years. This update isn’t just incremental; it’s a strategic modernization of compute, storage, networking, security, and operations across the entire private cloud stack.
.
Let’s break down the biggest enhancements and why they I think they matter.
Modernizing Infrastructure Economics with vSphere Foundation 9.1
VCF 9.1 brings several powerful updates to the vSphere layer, aimed at improving performance efficiency and reducing operational overhead.
Enhanced NVMe Memory Tiering
Workloads that demand high throughput and low latency benefit from smarter memory tiering. NVMe-based memory tiers now deliver improved performance and flexibility. (And yes — many are hoping Secure Boot support lands here as well.)
Parallel Processing of DRS vMotion
DRS can now process multiple vMotions in parallel, dramatically reducing cluster balancing times. This is especially impactful in large-scale environments with frequent workload mobility.
Live Patching for TPM-Enabled Hosts
Live patching now works even on hosts with TPM enabled — a huge win for security-conscious organizations that previously had to choose between uptime and compliance.
Networking Updates: Scale, Simplicity, and Smarter Automation
VCF 9.1 introduces major networking enhancements that streamline operations and expand connectivity options.
Enhanced Day-2 VM Lifecycle Management
Networking changes for VMs — including NIC updates, IP changes, and security policies — are now easier and more automated.
Existing VLAN Connectivity via Distributed Transit Gateways
You can now bridge existing VLAN-based networks into VCF environments more seamlessly, reducing migration friction and simplifying hybrid designs.
VCF 9.1 now supports EVPN-VXLAN interoperability with the physicalnetwork fabric. This is a major step toward fully integrated, fabric-aware cloud networking.
Network Assessment & VPC Planning
New tools and workflows help architects plan VPC layouts, assess network readiness, and avoid misconfigurations before deployment.
Optimize, Modernize & Protect Storage with vSAN in VCF 9.1
Storage gets a significant upgrade in this release, especially for environments focused on efficiency and resilience.
Encryption for vSAN Global Deduplication
Global dedupe is now compatible with data-at-rest encryption — a long-awaited capability for secure, space-efficient storage.
Enhanced Stretched Cluster Capabilities
Improved resilience and smarter failure handling strengthen business continuity for mission-critical workloads.
Automated Storage Policy Management
Policies now adjust automatically based on cluster configuration changes, reducing manual tuning and risk of misalignment.
Strengthening Zero Trust Security & Platform Resilience
Security is a major theme in VCF 9.1, with improvements across the stack.
Data-at-Rest Encryption for Global Dedupe
This ensures encrypted storage without sacrificing dedupe efficiency — a rare combination in enterprise storage.
Quick Patching for vCenter
Faster patch cycles reduce exposure windows and simplify maintenance.
Live Patching for TPM-Enabled Hosts
As mentioned earlier, this is a major operational win for secure environments.
Continuous Compliance & Integrated Cyber Recovery
VCF 9.1 pushes deeper into automated compliance and recovery workflows.
Compliance Monitoring & Desired State Remediation
The platform now continuously checks VCF components against desired state and can automatically remediate drift.
VPC Policy-Based Connectivity
Security and connectivity policies can now be applied consistently across VPCs, improving governance and reducing misconfigurations.
VMware Data Services Manager 9.1: Modern Databases for AI & Cloud
Microsoft SQL Server 2022 Now GA
SQL Server 2022 is now fully supported and generally available through DSM 9.1, enabling automated lifecycle management for modern database workloads — including those powering AI and analytics.
Want to See It in Action?
VMware has published a full VCF 9.1 video podcast series that dives deeper into the new capabilities:
Enough to do in my Homelab Starting with Upgrade and testing the new features!!
At first you have to deploy the VMware Cloud Foundation
The name of the installer is a little off: VCF-SDDC-Manager-Appliance-9.0.0.0.24703748.ova
Download token
If you don’t have a download token from the Broadcom support then it’s a little complicated. I won’t go into this in depth now but here is a nice article if you have a Synology Nas: VCF 9.0 Offline Depot using Synology
(Note: For any vExpert reading this… You need to claim your own domain name because build your profile is not working with a @hotmail.com or @gmail accounts, Noted by Broadcom)
I have deployed the beta in a earlier state. For deployment I used the same json file to start with.
VCF Automation will deployed automatically that was only the change that I changed it
It’s running!
Performance running Nested!!
Above the screenshot from the host with de Nested VCF lab deployment it’s quite cpu intensive (2 x CPU 8cores / 16 Threads)
Optimalization
After the deployment I did three things
– On the vSAN ESA Cluster (Disable Auto Disk Claim) (Keep the warning away)
– On the NSX VM reduce the cpu reservation from 6000 🡪 2000 (It helps but not enough)
– VCF automation VM 24 cores to 16 cores.
NSX en VCF automation are really CPU intensive.
I had deployed also two Edge servers but the server did not like that. Edges are also CPU intensive.
I am thinking about adding MS-01 or MS-A2 for splitting some load.
Websites store cookies to enhance functionality and personalise your experience. You can manage your preferences, but blocking some cookies may impact site performance and services.
Essential cookies enable basic functions and are necessary for the proper function of the website.
Name
Description
Duration
Cookie Preferences
This cookie is used to store the user's cookie consent preferences.
30 days
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Contains information related to marketing campaigns of the user. These are shared with Google AdWords / Google Ads when the Google Ads and Google Analytics accounts are linked together.
90 days
__utma
ID used to identify users and sessions
2 years after last activity
__utmt
Used to monitor number of Google Analytics server requests
10 minutes
__utmb
Used to distinguish new sessions and visits. This cookie is set when the GA.js javascript library is loaded and there is no existing __utmb cookie. The cookie is updated every time data is sent to the Google Analytics server.
30 minutes after last activity
__utmc
Used only with old Urchin versions of Google Analytics and not with GA.js. Was used to distinguish between new sessions and visits at the end of a session.
End of session (browser)
__utmz
Contains information about the traffic source or campaign that directed user to the website. The cookie is set when the GA.js javascript is loaded and updated when data is sent to the Google Anaytics server
6 months after last activity
__utmv
Contains custom information set by the web developer via the _setCustomVar method in Google Analytics. This cookie is updated every time new data is sent to the Google Analytics server.
2 years after last activity
__utmx
Used to determine whether a user is included in an A / B or Multivariate test.
18 months
_ga
ID used to identify users
2 years
_gali
Used by Google Analytics to determine which links on a page are being clicked
30 seconds
_ga_
ID used to identify users
2 years
_gid
ID used to identify users for 24 hours after last activity
24 hours
_gat
Used to monitor number of Google Analytics server requests when using Google Tag Manager
You must be logged in to post a comment.