Skip to the essay
ShemolSedna joint inference and federated learning controller optimization v1.1
云原生 / KubeEdge

Sedna joint inference and federated learning controller optimization v1.1

When federated learning implements changing config by editing the custom resource (kubectl edit FederatedLearningJob **), the update path is that a pod cannot have its config parameters updated in place the way a joint-inference custom resource's Deployment can (see this issue and this post). So the only option was to delete all the pods (because there was no way to get a single pod) and create them again with the new config. In the joint-inference controller, Deployment names are fixed at creation time, so you can

go
Deployment, err := c.deploymentsLister.Deployments(service.Namespace).Get(workerName)

get that specific Deployment, change parameters, and update. But right now generated pod names have 5 random characters at the end, so you can't look a pod up by name. I also found that the naming in the controller doesn't actually take effect (this code):

go
"WORKER_NAME": "aggworker-" + utilrand.String(5)

This line doesn't actually do anything, so Kubernetes is auto-naming the pod.

The reason is that injectWorkerParam in pkg/globalmanager/runtime/worker.go never assigns pod.ObjectMeta.Name. I think it should be

go
pod.ObjectMeta.Name = workerParam.Env["WORKER_NAME"]

Then you can pin the pod name by changing "WORKER_NAME". Once the name is known, and you know which worker field on the custom resource changed, you can delete that one pod and recreate it, instead of deleting everything.

Later Tang Ming pointed out that unlike inference jobs, a federated-learning training job is basically a Kubernetes Job. A Job is one-shot, so maybe just forbid in-place edits: if you want new parameters, delete and redeploy. I think that's reasonable. On top of that, pods themselves don't support updates, so this approach feels a bit against Kubernetes' intent... So the next optimization is actually restricting access to the resource. Still gathering material...

Compared with the Open Source Promotion Plan work, these are small changes, so this is v1.1. Next will be v2 because the thinking diverges from v1. So this won't be a PR either.