Download keadm
Download keadm to install KubeEdge. Official docs: https://kubeedge.io/docs/setup/install-with-keadm/
(the English version has the download section, the Chinese one doesn't — a bit confusing...)
wget <https://github.com/kubeedge/kubeedge/releases/download/v1.16.2/keadm-v1.16.2-linux-amd64.tar.gz>
tar -zxvf keadm-v1.16.2-linux-amd64.tar.gz
cp keadm-1.16.2-linux-amd64/keadm/keadm /usr/local/bin/keadm
Set up the cloud (KubeEdge master)
sudo keadm init --advertise-address=主机ip地址 --kubeedge-version=v1.16.2 --set iptablesManager.mode="external" --set cloudCore.modules.dynamicController.enable=true
There's also a flag --kube-config=/root/.kube/config, but that's the default so I dropped it. If your config file isn't at the default path you need to set it; you can also check with keadm --help.
If it doesn't succeed, there can be a million reasons (pain).
One reason might be you didn't remove the taint on the node, so you can't deploy workloads on it. Fix:
kubectl taint nodes master node-role.kubernetes.io/control-plane:NoSchedule-
Another reason might be the cloudcore image didn't pull. Then check whether the proxy you set for containerd actually took effect, or swap in a domestic image. I used 渡渡鸟镜像同步站 at the time; search and you can see the cloudcore v1.16.2 I used.
kubectl edit pod/daemonset/deployment ** -n **
Then just replace it with the domestic image address.
I also ran into disk pressure... but that should be pretty rare.
Anyway, when something goes wrong, just describe pod/node and look at the logs.
Once cloudcore is running normally, we keadm gettoken --kube-config=... to get the token, and get ready to deploy edgecore.
Set up the edge (KubeEdge worker)
On the edge device (containerd and keadm already installed):
sudo keadm join --kubeedge-version=1.16.2 --cloudcore-ipport="云端ip地址":10000 --remote-runtime-endpoint=unix:///run/containerd/containerd.sock --cgroupdriver=systemd --token=**
If you did everything above, this usually just works. If not, look at the logs, and search Github issues.
Deploy EdgeMesh
This is where I made the most mistakes, and where I spent the most time. If this isn't deployed right, the joint-inference example simply won't run.
Official docs: https://edgemesh.netlify.app/guide/
Enable the edge Kube-API endpoint
-
- Enable the dynamicController module on the cloud
If you deployed cloudcore with the command above, this is already on, because I added -set cloudCore.modules.dynamicController.enable=true.
-
- Turn on the metaServer module on the edge. After you finish the config, you have to restart edgecore.
vim /etc/kubeedge/config/edgecore.yaml
modules:
...
edgeMesh:
enable: false
...
metaManager:
metaServer:
enable: true
Restart edgecore
systemctl restart edgecore
-
- On the edge node, set clusterDNS and clusterDomain. After the config, you need to restart edgecore.
$ vim /etc/kubeedge/config/edgecore.yaml
modules:
...
edged:
...
tailoredKubeletConfig:
...
clusterDNS:
- 169.254.96.16
clusterDomain: cluster.local
...
Pay close attention to where you put the config (tears were shed)...
Don't change the clusterDNS value.
As the docs say, the clusterDNS value '169.254.96.16' comes from the default of bridgeDeviceIP in commonConfig在新窗口打开. Normally you don't need to change it; if you really must, keep the two in sync.
Restart edgecore.
systemctl restart edgecore
-
- Finally, on the edge node, test whether the edge Kube-API endpoint works:
$ curl 127.0.0.1:10550/api/v1/services
{"apiVersion":"v1","items":[{"apiVersion":"v1","kind":"Service","metadata":{"creationTimestamp":"2021-04-14T06:30:05Z","labels":{"component":"apiserver","provider":"kubernetes"},"name":"kubernetes","namespace":"default","resourceVersion":"147","selfLink":"default/services/kubernetes","uid":"55eeebea-08cf-4d1a-8b04-e85f8ae112a9"},"spec":{"clusterIP":"10.96.0.1","ports":[{"name":"https","port":443,"protocol":"TCP","targetPort":6443}],"sessionAffinity":"None","type":"ClusterIP"},"status":{"loadBalancer":{}}},{"apiVersion":"v1","kind":"Service","metadata":{"annotations":{"prometheus.io/port":"9153","prometheus.io/scrape":"true"},"creationTimestamp":"2021-04-14T06:30:07Z","labels":{"k8s-app":"kube-dns","kubernetes.io/cluster-service":"true","kubernetes.io/name":"KubeDNS"},"name":"kube-dns","namespace":"kube-system","resourceVersion":"203","selfLink":"kube-system/services/kube-dns","uid":"c221ac20-cbfa-406b-812a-c44b9d82d6dc"},"spec":{"clusterIP":"10.96.0.10","ports":[{"name":"dns","port":53,"protocol":"UDP","targetPort":53},{"name":"dns-tcp","port":53,"protocol":"TCP","targetPort":53},{"name":"metrics","port":9153,"protocol":"TCP","targetPort":9153}],"selector":{"k8s-app":"kube-dns"},"sessionAffinity":"None","type":"ClusterIP"},"status":{"loadBalancer":{}}}],"kind":"ServiceList","metadata":{"resourceVersion":"377360","selfLink":"/api/v1/services"}}
If the response is an empty list, or it takes a long time (close to 10s) to come back, your config is probably wrong. Check carefully.
After the steps above, KubeEdge's edge Kube-API endpoint is on. Then you can keep deploying EdgeMesh.
Start deploying EdgeMesh
Follow the docs' prerequisites: clear taints, add a filter label.
$ kubectl taint nodes --all node-role.kubernetes.io/master-
$ kubectl label services kubernetes service.edgemesh.kubeedge.io/service-proxy-name=""
Then install EdgeMesh by hand:
- Clone EdgeMesh from Github
$ git clone <https://github.com/kubeedge/edgemesh.git>
$ cd edgemesh
- Create the CRDs
$ kubectl apply -f build/crds/istio/
customresourcedefinition.apiextensions.k8s.io/destinationrules.networking.istio.io created
customresourcedefinition.apiextensions.k8s.io/gateways.networking.istio.io created
customresourcedefinition.apiextensions.k8s.io/virtualservices.networking.istio.io created
- Deploy edgemesh-agent
There's a note below: you need to edit the relayNodes part of build/agent/resources/04-configmap.yaml, and regenerate the PSK.
For relaynode you usually pick one cloud node as the relay — the master, and the ip is the master's ip. Generate the PSK from the URL in the comments. For your own experiments, not generating it is also fine (bushi
relayNodes:
- nodeName: cloud #master的名字
advertiseAddress:
- *.*.*.* #master的ip
Comment out the rest.
Then deploy edgemesh-agent
$ kubectl apply -f build/agent/resources/
serviceaccount/edgemesh-agent created
clusterrole.rbac.authorization.k8s.io/edgemesh-agent created
clusterrolebinding.rbac.authorization.k8s.io/edgemesh-agent created
configmap/edgemesh-agent-cfg created
configmap/edgemesh-agent-psk created
daemonset.apps/edgemesh-agent created
- Check it
$ kubectl get all -n kubeedge -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
pod/edgemesh-agent-7gf7g 1/1 Running 0 39s 192.168.0.71 k8s-node1 <none> <none>
pod/edgemesh-agent-fwf86 1/1 Running 0 39s 192.168.0.229 k8s-master <none> <none>
pod/edgemesh-agent-twm6m 1/1 Running 0 39s 192.168.5.121 ke-edge2 <none> <none>
pod/edgemesh-agent-xwxlp 1/1 Running 0 39s 192.168.5.187 ke-edge1 <none> <none>
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE CONTAINERS IMAGES SELECTOR
daemonset.apps/edgemesh-agent 4 4 4 4 4 <none> 39s edgemesh-agent kubeedge/edgemesh-agent:latest k8s-app=kubeedge,kubeedge=edgemesh-agent
This is the official check. I don't find it that useful — even running can still mean it isn't actually working.
On the edge you can
crictl logs edgemesh的containerID
and check whether the logs look normal, or where it's stuck. If it's healthy you'll see heartbeat sent on a schedule.
Run the EdgeMesh test case
Strongly recommend running it — you'll find a lot of problems. When I asked seniors for help, they always asked first whether the test case passed. I ran the starred one, Cross-Edge-Cloud.
-
- Deploy the test pods.
$ kubectl apply -f examples/test-pod.yaml
pod/alpine-test created
pod/websocket-test created
-
- Deploy what the edge–cloud comms test needs
$ kubectl apply -f examples/cloudzone.yaml
namespace/cloudzone created
deployment.apps/tcp-echo-cloud created
service/tcp-echo-cloud-svc created
deployment.apps/busybox-sleep-cloud created
$ kubectl apply -f examples/edgezone.yaml
namespace/edgezone created
deployment.apps/tcp-echo-edge created
service/tcp-echo-edge-svc created
deployment.apps/busybox-sleep-edge created
When the experiment got to this step I hit the first problem: the pod that should land on the edge, namespace edgezone, stayed in ContainerCreating. Because the Pod wouldn't deploy, of course I couldn't look at logs (crictl logs containerID), couldn't look from the cloud either (kubectl logs), and kubectl describe pod had zero info. So I did systemctl status edgecore, and finally saw the related error (though maybe there's a better way?). The reason was /run/flannel/subnet.env didn't exist on the edge. I checked and the cloud had it, so I created a file on the edge with the same contents. After a little while the pods all showed running.
-
- Cloud accessing edge
$ BUSYBOX_POD=$(kubectl get all -n cloudzone | grep pod/busybox | awk '{print $1}')
$ kubectl -n cloudzone exec $BUSYBOX_POD -c busybox -i -t -- sh
$ telnet tcp-echo-edge-svc.edgezone 2701
Welcome, you are connected to node ke-edge1.
Running on Pod tcp-echo-edge.
In namespace edgezone.
With IP address 172.17.0.2.
Service default.
Hello Edge, I am Cloud.
Hello Edge, I am Cloud.
I didn't hit any problem here.
-
- Edge accessing cloud
The official site uses docker, so they used docker ps. With crictl just
crictl ps
# 找到busybox的containerID,然后
crictl exec -it containerID sh
Then I ran
$ telnet tcp-echo-cloud-svc.cloudzone 2701
and hit a problem, something like name or server unknow. That was actually because I'd misconfigured the edge Kube-API.
After I fixed the config and tried this step again, another problem: no route to host, same as this issue. Then I followed question 3 in the most complete EdgeMesh Q&A handbook mentioned in the issue, cleaned iptables rules, redeployed edgemesh, and it went through! I was so hyped at the time.
$ telnet tcp-echo-cloud-svc.cloudzone 2701
Welcome, you are connected to node k8s-master.
Running on Pod tcp-echo-cloud.
In namespace cloudzone.
With IP address 10.244.0.8.
Service default.
Hello Cloud, I am Edge.
Hello Cloud, I am Edge.
That's when EdgeMesh was actually deployed successfully. I don't even know how much time I spent on EdgeMesh. But that's just what learning takes. At first I had no habit of looking at logs, I just randomly tried things. Now when something breaks, the first reaction is to find the logs, and then you can fix it fast.
Deploy Sedna
Official docs: https://sedna.readthedocs.io/en/latest/setup/install.html
curl <https://raw.githubusercontent.com/kubeedge/sedna/main/scripts/installation/install.sh> | SEDNA_ACTION=create bash -
One issue: this script sometimes fails to detect the version, so watch what it prints. If it doesn't pick up a version, interrupt the install, uninstall Sedna, and try again.
# 卸载的命令
curl <https://raw.githubusercontent.com/kubeedge/sedna/main/scripts/installation/install.sh> | SEDNA_ACTION=create bash -
If it's healthy, it just runs. If not, you may still need to swap in a domestic image.
References
- Full k8s+kubeedge+sedna install + pitfalls + fixes: https://blog.csdn.net/MacWx/article/details/130200209