Skip to the essay
ShemolInstalling a k8s Cluster and Preparing to Deploy KubeEdge
流程记录 / 云原生

Installing a k8s Cluster and Preparing to Deploy KubeEdge

Lab environment: Debian GNU/Linux 12 (bookworm) x86_64

Install containerd

Before installing k8s, you need to install containerd first. The KubeEdge edge side only needs containerd; the cloud side needs k8s.

Prep before installing k8s

Enable traffic forwarding

Tweak the iptables config and enable the "br_netfilter" module, so Kubernetes can inspect and forward network traffic.

shell
cat <<EOF | sudo tee /etc/modules-load.d/k8s.conf
br_netfilter
EOF

cat <<EOF | sudo tee /etc/sysctl.d/k8s.conf
net.bridge.bridge-nf-call-ip6tables = 1
net.bridge.bridge-nf-call-iptables = 1
net.ipv4.ip_forward=1 # better than modify /etc/sysctl.conf
EOF

sudo sysctl --system

Turn off Linux Swap

For security reasons (the official docs promise Secrets are only read and written in memory, never persisted to disk), and to keep node sync consistent, starting from 1.8 Kubernetes explicitly says it doesn't support Swap by default. On a machine where Swap isn't turned off, the cluster simply won't start.
shell
sudo cp /etc/fstab /etc/fstab_bak
sudo swapoff -a
sudo sed -ri '/\\sswap\\s/s/^#?/#/' /etc/fstab

Register the apt repo

I went with the Tsinghua mirror. Official docs: https://mirrors.tuna.tsinghua.edu.cn/help/kubernetes/

shell
sudo apt install -y apt-transport-https ca-certificates curl

sudo curl -fsSLo /usr/share/keyrings/kubernetes-archive-keyring.gpg <https://packages.cloud.google.com/apt/doc/apt-key.gpg>

Create /etc/apt/sources.list.d/kubernetes.list, with this content:

shell
deb [signed-by=/usr/share/keyrings/kubernetes-archive-keyring.gpg] <https://mirrors.tuna.tsinghua.edu.cn/kubernetes/apt> kubernetes-xenial main

Then

shell
sudo apt update

Install kubeadm, kubelet, kubectl

Please look up what these three tools do yourself.

shell
sudo apt install kubeadm kubelet kubectl

I didn't pin versions. If you need to pin versions, please read the reference article.

Check the install:

shell
cloud@cloud:~$ kubeadm version
kubeadm version: &version.Info{Major:"1", Minor:"28", GitVersion:"v1.28.10", GitCommit:"21be1d76a90bc00e2b0f6676a664bdf097224155", GitTreeState:"clean", BuildDate:"2024-05-14T10:51:30Z", GoVersion:"go1.21.9", Compiler:"gc", Platform:"linux/amd64"}
cloud@cloud:~$ kubectl version
Client Version: v1.28.10
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
The connection to the server localhost:8080 was refused - did you specify the right host or port?
cloud@cloud:~$ kubelet --version
Kubernetes v1.28.10

See which image versions you'll need later:

shell
cloud@cloud:~$ sudo kubeadm config images list --kubernetes-version v1.28.10
registry.k8s.io/kube-apiserver:v1.28.10
registry.k8s.io/kube-controller-manager:v1.28.10
registry.k8s.io/kube-scheduler:v1.28.10
registry.k8s.io/kube-proxy:v1.28.10
registry.k8s.io/pause:3.9
registry.k8s.io/etcd:3.5.12-0
registry.k8s.io/coredns/coredns:v1.10.1

Init the cluster control plane

  1. Start kubelet, and make sure it runs on boot
shell
sudo systemctl start kubelet
sudo systemctl enable kubelet
  1. Start deploying. Here I switched to the root account
shell
kubeadm init \\
--image-repository registry.cn-hangzhou.aliyuncs.com/google_containers \\
--pod-network-cidr=10.10.0.0/16 \\
--apiserver-advertise-address=10.129.196.8 \\
--kubernetes-version=v1.28.10 \\
--v=5

Flag explanations (yeah, I just copied these from the original reference article):

  • -image-repository: pull the base images above from Aliyun; if you don't set this, you have to pull from Google's servers;
  • -pod-network-cidr: set the Pod CIDR in the cluster; this is for installing the Flannel network plugin later;
  • -kubernetes-version: specify the Kubernetes version;
  • -v=5: print detailed trace logs, see here;
  • -apiserver-advertise-address: specify the api-server IP. If you have multiple NICs, pick which one explicitly. apiserver is a big deal in a Kubernetes cluster — lots of config (ConfigMaps and so on) store that address directly, and changing it later is a huge pain, so be careful.
  1. Follow the log prompts to set up kube config

My approach was to do this as root. Later when deploying kubeedge, the default config file is at $HOME/.kube/config, which is /root/.kube/config.

shell
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config

Install the network plugin

Please learn about k8s network plugins yourself — I still need to study this more too...

Here we pick Flannel, same as the original reference article.

Get the kube-flannel.yml install file from Github, then edit it:

plain text
net-conf.json: |
    {
      "Network": "10.244.0.0/16",
      "Backend": {
        "Type": "vxlan"
      }
    }

Change the Network above to the pod CIDR we set when initializing the cluster:

yaml
net-conf.json: |
    {
      "Network": "10.10.0.0/16",
      "Backend": {
        "Type": "vxlan"
      }
    }

Finally install with kubectl apply:

plain text
kubectl apply -f kube-flannel.yml

You can also do it the flannel official way:

shell
kubectl apply -f <https://github.com/flannel-io/flannel/releases/latest/download/kube-flannel.yml>

Version-wise I feel either is probably fine, but either way the pod CIDR should be configured.

One small thing to watch for

Later when running the edgemesh test case, I found that pods deployed on the edge stayed in ContainerCreating. The edge edgecore logs said /run/flannel/subnet.env didn't exist. After I copied that file's contents from the cloud side to the edge, they deployed normally. So keep an eye on this; I'll mention it again when deploying edgemesh.

About images

Before I set up a proxy for containerd, I ran into flannel images that wouldn't pull. Here's a domestic mirror workaround I used.

Search for the relevant images on the 渡渡鸟镜像同步站, then kubectl edit pod/deployment/daemonset ** -n ** and swap in the domestic image. If 渡渡鸟 doesn't have the image, you can add it following the site's usage guide.

Remove the taint on master

Removing the taint is so the master node can run actual workloads. When I was deploying cloudcore this issue cost me some time...

shell
kubectl taint nodes cloud node-role.kubernetes.io/control-plane:NoSchedule-

The reference article also talks about adjusting the NodePort range. I haven't hit that yet, so I skipped it — read the reference if you need it.

Adding Worker nodes isn't necessary for now either; read the reference if you need it.

Install Docker

Also, when running the joint inference example, you need to build images for the helmet-detection large/small models (run the build_image.sh script in the example — I'll talk about that later), so you need docker ce.

Install docker ce from the Tsinghua mirror. Official docs: https://mirrors.tuna.tsinghua.edu.cn/help/docker-ce/

Docker should also get a proxy and a domestic mirror. Search for how to do that yourself.

References