On September 3, Broadcom introduced VMware AI Factory as part of VMware Private AI Cloud. The software prepares servers to run models and helps maintain them after deployment. The vendor promises to cut the journey from hardware setup to the first running model from weeks to hours.

The model physically runs on a server, a computer employees access over a network. A local assistant also needs an application that accepts questions, searches documents and displays an answer. VMware automates preparation of the computing environment for that software.

One model for several departments

In VCF 9.1.1, Broadcom added model sharing with separate departmental data. One model-runtime service can serve the organisation while each department’s data stays in its own access area.

With separate installations, each department loads its own model copy into accelerator memory. A shared service allows different applications to use one model. The documents do not become shared: access separation applies to the data those applications use.

Preparing servers from a shared console

AI Factory relies on VMware Cloud Foundation, or VCF, a platform for managing compute, storage and networking. The announcement names compatible Dell, Cisco, Lenovo and Supermicro servers, AMD Instinct accelerators and a partnership with MetalSoft to automate physical hardware preparation.

MetalSoft is intended to automate initial physical-server preparation and firmware operations. Administrators will be able to manage hardware from different vendors through the VCF console rather than use each supplier’s tools separately.

What the promised hours leave out

ITPro reports the shorter deployment time as a Broadcom claim, not the result of an independent comparison between two identical installations.

The promised deployment time ends with the first running model. Connecting an internal archive is a separate part of the application: extracting document text, arranging search and linking retrieved passages to an answer. Installing server software does not itself perform that work.

What comes later

In its future-release plans, Broadcom lists AI Gateway, which will select local or cloud models for requests and limit token use. The same section includes isolation for software that agents generate and execute.

Model autoscaling is described separately. Administrators set response-time and session-count thresholds; the system should allocate more resources as demand rises and release them as it falls. Computing resources could then move between tasks without continual manual reassignment.

Unlike model sharing in VCF 9.1.1, these capabilities remain assigned to future releases. Broadcom’s announcement does not specify a version or date for model autoscaling to become available.