Update the cluster management and advanced deployment docs (#10790)

This commit is contained in:
吴晟 Wu Sheng 2023-05-11 15:56:48 +08:00 committed by GitHub
parent 54ae820fb2
commit 7a0331a319
No known key found for this signature in database
GPG Key ID: 4AEE18F83AFDEB23
3 changed files with 54 additions and 47 deletions

View File

@ -71,5 +71,6 @@
* Add Profiling related documentations.
* Add `SUM_PER_MIN` to MAL documentation.
* Make the log relative docs more clear, and easier for further more formats support.
* Update the cluster management and advanced deployment docs.
All issues and pull requests are [here](https://github.com/apache/skywalking/milestone/169?closed=1)

View File

@ -1,6 +1,8 @@
# Advanced deployment
OAP servers communicate with each other in a cluster environment.
In the cluster mode, you could run in different roles.
OAP servers communicate with each other in a cluster environment to do distributed aggregation.
In the cluster mode, all OAP nodes are running in Mixed mode by default.
The available roles for OAP are,
- Mixed(default)
- Receiver
- Aggregator
@ -24,7 +26,7 @@ The OAP is responsible for:
## Aggregator
The OAP is responsible for:
1. Internal communication(receive)
1. Internal communication(receiving from Receiver and Mixed roles OAP)
1. L2 aggregation
1. Persistence
1. Alarm
@ -33,5 +35,7 @@ ___
These roles are designed for complex deployment requirements on security and network policy.
## Kubernetes
If you are using our native [Kubernetes coordinator](backend-cluster.md#kubernetes), the `labelSelector`
setting is used for `Aggregator` role selection rules. Choose the right OAP deployment based on your needs.
If you are using our native [Kubernetes coordinator](backend-cluster.md#kubernetes), and you insist to install OAP nodes
with a clearly defined role. There should be two deployments for each role,
one for receiver OAPs and the other for aggregator OAPs to separate different system environment settings.
Then, the `labelSelector` should be set for `Aggregator` role selection rules to choose the right OAP deployment based on your needs.

View File

@ -1,28 +1,59 @@
# Cluster Management
In many production environments, the backend needs to support high throughput and provide high availability (HA) to
maintain robustness,
so you always need cluster management in product env.
In many production environments, the backend needs to support **distributed aggregation**, high throughput
and provide high availability (HA) to maintain robustness, so **you always need to setup CLUSTER management in product env**.
Otherwise, you would face metrics **inaccurate**.
`core/gRPCHost` is listening on `0.0.0.0` for quick start as the single mode for most cases.
Besides the `Kubernetes` coordinator, which is using the cloud-native mode to establish cluster, all other coordinators
requires `core/gRPCHost` updated to real IP addresses or take reference of `internalComHost` and `internalComPort` in each
coordinator doc.
NOTICE, cluster management doesn't provide a service discovery mechanism for agents and probes. We recommend
agents/probes using
gateway to load balancer to access OAP clusters.
The core feature of cluster management is supporting the whole OAP cluster running distributed aggregation and analysis
for telemetry data.
agents/probes using gateway to load balancer to access OAP clusters.
There are various ways to manage the cluster in the backend. Choose the one that best suits your needs.
- [Zookeeper coordinator](#zookeeper-coordinator). Use Zookeeper to let the backend instances detect and communicate
with each other.
- [Kubernetes](#kubernetes). When the backend clusters are deployed inside Kubernetes, you could make use of this method
by using k8s native APIs to manage clusters.
- [Zookeeper coordinator](#zookeeper-coordinator). Use Zookeeper to let the backend instances detect and communicate
with each other.
- [Consul](#consul). Use Consul as the backend cluster management implementor and coordinate backend instances.
- [Etcd](#etcd). Use Etcd to coordinate backend instances.
- [Nacos](#nacos). Use Nacos to coordinate backend instances.
In the `application.yml` file, there are default configurations for the aforementioned coordinators under the
section `cluster`.
You can specify any of them in the `selector` property to enable it.
In the `application.yml` file, there are default configurations for the aforementioned coordinators under the
section `cluster`. You can specify any of them in the `selector` property to enable it.
## Kubernetes
The required backend clusters are deployed inside Kubernetes. See the guides in [Deploy in kubernetes](backend-k8s.md).
Set the selector to `kubernetes`.
```yaml
cluster:
selector: ${SW_CLUSTER:kubernetes}
# other configurations
```
Meanwhile, the OAP cluster requires the pod's UID which is laid at `metadata.uid` as the value of the system environment variable **SKYWALKING_COLLECTOR_UID**
```yaml
containers:
# Original configurations of OAP container
- name: {{ .Values.oap.name }}
image: {{ .Values.oap.image.repository }}:{{ required "oap.image.tag is required" .Values.oap.image.tag }}
# ...
# ...
env:
# Add metadata.uid as the system environment variable, SKYWALKING_COLLECTOR_UID
- name: SKYWALKING_COLLECTOR_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
```
Read [the complete helm](https://github.com/apache/skywalking-kubernetes/blob/476afd51d44589c77a4cbaac950272cd5d064ea9/chart/skywalking/templates/oap-deployment.yaml#L125) for more details.
## Zookeeper coordinator
@ -73,36 +104,7 @@ zookeeper:
enableACL: ${SW_ZK_ENABLE_ACL:false} # disable ACL in default
schema: ${SW_ZK_SCHEMA:digest} # only support digest schema
expression: ${SW_ZK_EXPRESSION:skywalking:skywalking}
```
## Kubernetes
The required backend clusters are deployed inside Kubernetes. See the guides in [Deploy in kubernetes](backend-k8s.md).
Set the selector to `kubernetes`.
```yaml
cluster:
selector: ${SW_CLUSTER:kubernetes}
# other configurations
```
Meanwhile, the OAP cluster requires the pod's UID which is laid at `metadata.uid` as the value of the system environment variable **SKYWALKING_COLLECTOR_UID**
```yaml
containers:
# Original configurations of OAP container
- name: {{ .Values.oap.name }}
image: {{ .Values.oap.image.repository }}:{{ required "oap.image.tag is required" .Values.oap.image.tag }}
# ...
# ...
env:
# Add metadata.uid as the system environment variable, SKYWALKING_COLLECTOR_UID
- name: SKYWALKING_COLLECTOR_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
```
Read [the complete helm](https://github.com/apache/skywalking-kubernetes/blob/476afd51d44589c77a4cbaac950272cd5d064ea9/chart/skywalking/templates/oap-deployment.yaml#L125) for more details.
## Consul