Skip to the essay
ShemolDeploying KubeEdge, EdgeMesh, and Sedna
流程记录 / 云原生

Deploying KubeEdge, EdgeMesh, and Sedna

Download keadm

Download keadm to install KubeEdge. Official docs: https://kubeedge.io/docs/setup/install-with-keadm/

(the English version has the download section, the Chinese one doesn't — a bit confusing...)

shell
wget <https://github.com/kubeedge/kubeedge/releases/download/v1.16.2/keadm-v1.16.2-linux-amd64.tar.gz>

tar -zxvf keadm-v1.16.2-linux-amd64.tar.gz
cp keadm-1.16.2-linux-amd64/keadm/keadm /usr/local/bin/keadm

Set up the cloud (KubeEdge master)

shell
sudo keadm init --advertise-address=主机ip地址 --kubeedge-version=v1.16.2 --set iptablesManager.mode="external" --set cloudCore.modules.dynamicController.enable=true

There's also a flag --kube-config=/root/.kube/config, but that's the default so I dropped it. If your config file isn't at the default path you need to set it; you can also check with keadm --help.

If it doesn't succeed, there can be a million reasons (pain).

One reason might be you didn't remove the taint on the node, so you can't deploy workloads on it. Fix:

shell
kubectl taint nodes master node-role.kubernetes.io/control-plane:NoSchedule-

Another reason might be the cloudcore image didn't pull. Then check whether the proxy you set for containerd actually took effect, or swap in a domestic image. I used 渡渡鸟镜像同步站 at the time; search and you can see the cloudcore v1.16.2 I used.

shell
kubectl edit pod/daemonset/deployment ** -n **

Then just replace it with the domestic image address.

I also ran into disk pressure... but that should be pretty rare.

Anyway, when something goes wrong, just describe pod/node and look at the logs.

Once cloudcore is running normally, we keadm gettoken --kube-config=... to get the token, and get ready to deploy edgecore.

Set up the edge (KubeEdge worker)

On the edge device (containerd and keadm already installed):

shell
sudo keadm join  --kubeedge-version=1.16.2 --cloudcore-ipport="云端ip地址":10000  --remote-runtime-endpoint=unix:///run/containerd/containerd.sock --cgroupdriver=systemd --token=**

If you did everything above, this usually just works. If not, look at the logs, and search Github issues.

Deploy EdgeMesh

This is where I made the most mistakes, and where I spent the most time. If this isn't deployed right, the joint-inference example simply won't run.

Official docs: https://edgemesh.netlify.app/guide/

Enable the edge Kube-API endpoint

    1. Enable the dynamicController module on the cloud

If you deployed cloudcore with the command above, this is already on, because I added -set cloudCore.modules.dynamicController.enable=true.

    1. Turn on the metaServer module on the edge. After you finish the config, you have to restart edgecore.
yaml
vim /etc/kubeedge/config/edgecore.yaml
modules:
  ...
  edgeMesh:
    enable: false
  ...
  metaManager:
    metaServer:
      enable: true

Restart edgecore

shell
systemctl restart edgecore
    1. On the edge node, set clusterDNS and clusterDomain. After the config, you need to restart edgecore.
plain text
$ vim /etc/kubeedge/config/edgecore.yaml
modules:
  ...
  edged:
    ...
    tailoredKubeletConfig:
      ...
      clusterDNS:
      - 169.254.96.16
      clusterDomain: cluster.local
...

Pay close attention to where you put the config (tears were shed)...

Don't change the clusterDNS value.

As the docs say, the clusterDNS value '169.254.96.16' comes from the default of bridgeDeviceIP in commonConfig在新窗口打开. Normally you don't need to change it; if you really must, keep the two in sync.

Restart edgecore.

shell
systemctl restart edgecore
    1. Finally, on the edge node, test whether the edge Kube-API endpoint works:
shell
$ curl 127.0.0.1:10550/api/v1/services
{"apiVersion":"v1","items":[{"apiVersion":"v1","kind":"Service","metadata":{"creationTimestamp":"2021-04-14T06:30:05Z","labels":{"component":"apiserver","provider":"kubernetes"},"name":"kubernetes","namespace":"default","resourceVersion":"147","selfLink":"default/services/kubernetes","uid":"55eeebea-08cf-4d1a-8b04-e85f8ae112a9"},"spec":{"clusterIP":"10.96.0.1","ports":[{"name":"https","port":443,"protocol":"TCP","targetPort":6443}],"sessionAffinity":"None","type":"ClusterIP"},"status":{"loadBalancer":{}}},{"apiVersion":"v1","kind":"Service","metadata":{"annotations":{"prometheus.io/port":"9153","prometheus.io/scrape":"true"},"creationTimestamp":"2021-04-14T06:30:07Z","labels":{"k8s-app":"kube-dns","kubernetes.io/cluster-service":"true","kubernetes.io/name":"KubeDNS"},"name":"kube-dns","namespace":"kube-system","resourceVersion":"203","selfLink":"kube-system/services/kube-dns","uid":"c221ac20-cbfa-406b-812a-c44b9d82d6dc"},"spec":{"clusterIP":"10.96.0.10","ports":[{"name":"dns","port":53,"protocol":"UDP","targetPort":53},{"name":"dns-tcp","port":53,"protocol":"TCP","targetPort":53},{"name":"metrics","port":9153,"protocol":"TCP","targetPort":9153}],"selector":{"k8s-app":"kube-dns"},"sessionAffinity":"None","type":"ClusterIP"},"status":{"loadBalancer":{}}}],"kind":"ServiceList","metadata":{"resourceVersion":"377360","selfLink":"/api/v1/services"}}

If the response is an empty list, or it takes a long time (close to 10s) to come back, your config is probably wrong. Check carefully.

After the steps above, KubeEdge's edge Kube-API endpoint is on. Then you can keep deploying EdgeMesh.

Start deploying EdgeMesh

Follow the docs' prerequisites: clear taints, add a filter label.

shell
$ kubectl taint nodes --all node-role.kubernetes.io/master-
$ kubectl label services kubernetes service.edgemesh.kubeedge.io/service-proxy-name=""

Then install EdgeMesh by hand:

  • Clone EdgeMesh from Github
shell
$ git clone <https://github.com/kubeedge/edgemesh.git>
$ cd edgemesh
  • Create the CRDs
plain text
$ kubectl apply -f build/crds/istio/
customresourcedefinition.apiextensions.k8s.io/destinationrules.networking.istio.io created
customresourcedefinition.apiextensions.k8s.io/gateways.networking.istio.io created
customresourcedefinition.apiextensions.k8s.io/virtualservices.networking.istio.io created
  • Deploy edgemesh-agent

There's a note below: you need to edit the relayNodes part of build/agent/resources/04-configmap.yaml, and regenerate the PSK.

For relaynode you usually pick one cloud node as the relay — the master, and the ip is the master's ip. Generate the PSK from the URL in the comments. For your own experiments, not generating it is also fine (bushi

yaml
relayNodes:
- nodeName: cloud #master的名字
  advertiseAddress:
  - *.*.*.*  #master的ip

Comment out the rest.

Then deploy edgemesh-agent

plain text
$ kubectl apply -f build/agent/resources/
serviceaccount/edgemesh-agent created
clusterrole.rbac.authorization.k8s.io/edgemesh-agent created
clusterrolebinding.rbac.authorization.k8s.io/edgemesh-agent created
configmap/edgemesh-agent-cfg created
configmap/edgemesh-agent-psk created
daemonset.apps/edgemesh-agent created
  • Check it
shell
$ kubectl get all -n kubeedge -o wide
NAME                       READY   STATUS    RESTARTS   AGE   IP              NODE         NOMINATED NODE   READINESS GATES
pod/edgemesh-agent-7gf7g   1/1     Running   0          39s   192.168.0.71    k8s-node1    <none>           <none>
pod/edgemesh-agent-fwf86   1/1     Running   0          39s   192.168.0.229   k8s-master   <none>           <none>
pod/edgemesh-agent-twm6m   1/1     Running   0          39s   192.168.5.121   ke-edge2     <none>           <none>
pod/edgemesh-agent-xwxlp   1/1     Running   0          39s   192.168.5.187   ke-edge1     <none>           <none>

NAME                            DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR   AGE   CONTAINERS       IMAGES                           SELECTOR
daemonset.apps/edgemesh-agent   4         4         4       4            4           <none>          39s   edgemesh-agent   kubeedge/edgemesh-agent:latest   k8s-app=kubeedge,kubeedge=edgemesh-agent

This is the official check. I don't find it that useful — even running can still mean it isn't actually working.

On the edge you can

shell
crictl logs edgemesh的containerID

and check whether the logs look normal, or where it's stuck. If it's healthy you'll see heartbeat sent on a schedule.

Run the EdgeMesh test case

Strongly recommend running it — you'll find a lot of problems. When I asked seniors for help, they always asked first whether the test case passed. I ran the starred one, Cross-Edge-Cloud.

    1. Deploy the test pods.
shell
$ kubectl apply -f examples/test-pod.yaml
pod/alpine-test created
pod/websocket-test created
    1. Deploy what the edge–cloud comms test needs
shell
$ kubectl apply -f examples/cloudzone.yaml
namespace/cloudzone created
deployment.apps/tcp-echo-cloud created
service/tcp-echo-cloud-svc created
deployment.apps/busybox-sleep-cloud created
shell
$ kubectl apply -f examples/edgezone.yaml
namespace/edgezone created
deployment.apps/tcp-echo-edge created
service/tcp-echo-edge-svc created
deployment.apps/busybox-sleep-edge created

When the experiment got to this step I hit the first problem: the pod that should land on the edge, namespace edgezone, stayed in ContainerCreating. Because the Pod wouldn't deploy, of course I couldn't look at logs (crictl logs containerID), couldn't look from the cloud either (kubectl logs), and kubectl describe pod had zero info. So I did systemctl status edgecore, and finally saw the related error (though maybe there's a better way?). The reason was /run/flannel/subnet.env didn't exist on the edge. I checked and the cloud had it, so I created a file on the edge with the same contents. After a little while the pods all showed running.

    1. Cloud accessing edge
plain text
$ BUSYBOX_POD=$(kubectl get all -n cloudzone | grep pod/busybox | awk '{print $1}')
$ kubectl -n cloudzone exec $BUSYBOX_POD -c busybox -i -t -- sh
$ telnet tcp-echo-edge-svc.edgezone 2701
Welcome, you are connected to node ke-edge1.
Running on Pod tcp-echo-edge.
In namespace edgezone.
With IP address 172.17.0.2.
Service default.
Hello Edge, I am Cloud.
Hello Edge, I am Cloud.

I didn't hit any problem here.

    1. Edge accessing cloud

The official site uses docker, so they used docker ps. With crictl just

shell
crictl ps
# 找到busybox的containerID,然后
crictl exec -it containerID sh

Then I ran

plain text
$ telnet tcp-echo-cloud-svc.cloudzone 2701

and hit a problem, something like name or server unknow. That was actually because I'd misconfigured the edge Kube-API.

After I fixed the config and tried this step again, another problem: no route to host, same as this issue. Then I followed question 3 in the most complete EdgeMesh Q&A handbook mentioned in the issue, cleaned iptables rules, redeployed edgemesh, and it went through! I was so hyped at the time.

plain text
$ telnet tcp-echo-cloud-svc.cloudzone 2701
Welcome, you are connected to node k8s-master.
Running on Pod tcp-echo-cloud.
In namespace cloudzone.
With IP address 10.244.0.8.
Service default.
Hello Cloud, I am Edge.
Hello Cloud, I am Edge.

That's when EdgeMesh was actually deployed successfully. I don't even know how much time I spent on EdgeMesh. But that's just what learning takes. At first I had no habit of looking at logs, I just randomly tried things. Now when something breaks, the first reaction is to find the logs, and then you can fix it fast.

Deploy Sedna

Official docs: https://sedna.readthedocs.io/en/latest/setup/install.html

shell
curl <https://raw.githubusercontent.com/kubeedge/sedna/main/scripts/installation/install.sh> | SEDNA_ACTION=create bash -

One issue: this script sometimes fails to detect the version, so watch what it prints. If it doesn't pick up a version, interrupt the install, uninstall Sedna, and try again.

shell
# 卸载的命令
curl <https://raw.githubusercontent.com/kubeedge/sedna/main/scripts/installation/install.sh> | SEDNA_ACTION=create bash -

If it's healthy, it just runs. If not, you may still need to swap in a domestic image.

References