Kafka

#126 · open · 21 comments

View on GitHub ↗

kun-song

## zero-copy 一般数据库处理查询时,首先将 disk 上的查询结果 **缓存** 在内存中,然后从缓存发进入网络传输,Kafka 通过 zero-copy 技术,可以实现直接将消息从文件(Linux 中是 filtesystem cache)发送到网络,避免在内存中进行拷贝,大大提升性能。 为啥能用 zero-cpoy 呢? 因为 disk 上保存的 message 与网络上传输的 message (producer -> broker -> consumer)的 **格式完全相同**,因此无需任何处理,直接从硬盘发送到网络即可。 ## 磁盘顺序读写 为持久化消息,Kafka 将消息保存在磁盘上,在我们的印象中,磁盘读写相对内存而言是非常慢的,实际上由于需要 **寻道**,磁盘的 **随机 IO** 非常慢,但 **顺序 IO** 要好很多。 Kafka 通过顺序读写磁盘提高了磁盘 IO 速度,起码达到可以接受的水平,大约几十兆每秒。 ## 分区 即使顺序 IO 能达到 100M/s 的读写速度,但依然很低,为实现水平扩展,Kafka 将 topic 分为 partition,每个 partition 互相独立,从而实现水平扩展。 比如有 100 个分区则可以实现 100M/s * 100 的吞吐量。 ## 压缩 抛开 Kafka,我们可以对单个消息进行压缩,但这种压缩带来的收益不大,因为大部分冗余是 **消息之间** 字段值重复导致的,同一类型的消息很可能具备一些重复字段,因此对消息进行批量压缩才能带来较大收益。 Kafka 支持对一批消息进行压缩,且 producer 将压缩后的消息发送到 broker,broker 将原样保存他们到硬盘时,只有这些压缩后的消息被 consumer 接受后处理时,才会进行解压,即 produer 压缩,broker 保持,consumer 解压。 ##

Comments

kun-song

## `session.timeout.ms` Kafka 文档定义: >The timeout used to detect consumer failures when using Kafka's group management facility. The consumer sends periodic heartbeats to indicate its liveness to the broker. If no heartbeats are received by the broker before the expiration of this session timeout, then the broker will remove this consumer from the group and initiate a rebalance. Note that the value must be in the allowable range as configured in the broker configuration by `group.min.session.timeout.ms` and `group.max.session.timeout.ms`. 首先,当 broker 认为有 consumer 崩溃(比如线程挂掉、网络异常等)时,broker 会把崩溃的 consumer 从 consumer group 中移除,并触发一次 rebalance。而 consumer 通过周期性地向 broker 发送心跳来表明自己还“活着”,如果 consumer 崩溃,则 broker 无法收到心跳。 问题是 broker 无法收到心跳并不一定是 consumer 崩溃导致的,暂时的网络拥塞、网络故障也可能导致 broker 无法收到心跳,这种故障可能几秒钟就自动恢复了,实际上,目前没有办法通过心跳区分 consumer 是不是真的挂掉了。 所以 Kafka 很粗暴的通过 `session.timeout.ms` 定义一个最大的“心跳空白时间”,即心跳正常表明 broker 和 consumer 之间存在正常会话(session),如果超过 `session.timeout.ms` 配置的时间 broker 没有收到心跳,则 broker 就认为 consumer 真的挂了。 注意,consumer 配置的会话超时时间受限于 broker 对应配置,即 `session.timeout.ms` 的值必须位于: * `group.min.session.timeout.ms` * `group.max.session.timeout.ms` 区间范围内。 ## `max.poll.interval.ms` [Notable changes in 0.10.1.0](https://kafka.apache.org/documentation/#upgrade_1010_notable): >The new Java Consumer now supports heartbeating from a background thread. There is a new configuration max.poll.interval.ms which controls the maximum time between poll invocations before the consumer will proactively leave the group (5 minutes by default). The value of the configuration request.timeout.ms must always be larger than max.poll.interval.ms because this is the maximum time that a JoinGroup request can block on the server while the consumer is rebalancing, so we have changed its default value to just above 5 minutes. Finally, the default value of session.timeout.ms has been adjusted down to 10 seconds, and the default value of max.poll.records has been changed to 500. https://stackoverflow.com/questions/39730126/difference-between-session-timeout-ms-and-max-poll-interval-ms-for-kafka-0-10-0 https://cwiki.apache.org/confluence/display/KAFKA/KIP-62%3A+Allow+consumer+to+send+heartbeats+from+a+background+thread https://markmail.org/message/oeg63goh3ed3qdap https://cwiki.apache.org/confluence/display/KAFKA/KIP-62%3A+Allow+consumer+to+send+heartbeats+from+a+background+thread

kun-song

# Kafka 会不会丢消息? The guarantee that Kafka offers is that a committed message will not be lost, as long as there is at least one in sync replica alive, at all times.

kun-song

# 优雅停机 https://kafka.apache.org/21/documentation/streams/developer-guide/write-streams.html

kun-song

https://stackoverflow.com/questions/51034806/messages-in-crashed-kafka-broker

kun-song

# KSQL https://docs.confluent.io/current/streams-ksql.html ## 安装 首先启动 KSQL 服务器: ``` bin/ksql-server-start -daemon etc/ksql/ksql-server.properties ``` 然后进入 KSQL: ``` ./ksql ``` ## 使用 ### 管理操作 查看当前 Kafka 集群的所有 topic 集合: ``` show topics; ``` 打印 topic 内容: ``` print 'topic-name' from beginning; ``` 删除流、表: ``` drop stream stream-name drop table table-name ``` 显示流、表的结构: ``` describe stream/table-name ``` ### 流操作 #### topic -> stream KSQL 无法自动推断消息各个字段的类型,因此需要明确指定: ``` create stream stream-name(sn VARCHAR, t BIGINT) with (KAFKA_TOPIC='topic-name', value_format='JSON'); ``` 其中 `value_format` 有 3 种取值: * `JSON` * `DELIMITED`(逗号分隔) * `AVRO`

kun-song

操作 | 代码 -- | -- 删除 stream | DROP STREAM click; 删除 table | DROP TABLE users; 显示stream/table的列名和类型 | DESCRIBE click; 显示stream/table的列名和类型以及kafka topic详细信息 | DESCRIBE EXTENDED click; 显示某个查询的执行计划 query_id通过show queries获取 | EXPLAIN query_id; 显示kafka topic内容(从头取数据) | PRINT 'click' FROM BEGINNING; 显示kafka topic内容(取最新数据) | PRINT 'click'; 打印topics | SHOW/LIST TOPICS; 打印streams | SHOW/LIST STREAMS; 打印tables | SHOW/LIST TABLES; 打印queries | SHOW QUERIES; 打印KSQL 配置 | SHOW PROPERTIES;

kun-song

ksql> show 24_m; line 1:6: no viable alternative at input 'show 24_m' Caused by: org.antlr.v4.runtime.NoViableAltException KSQL 基于 ANTLR

kun-song

https://yasina.me/detail/KSQL-syntax

kun-song

ksql默认是从kafka最新的数据查询消费的,如果你想从开头查询,则需要在会话上进行设置:SET 'auto.offset.reset' = 'earliest';

kun-song

# Replicated Logs: Quorums, ISRs, and State Machines (Oh my!) ISR min.insync.replicas 集群有 2f + 1 个 broker 节点,若: * producer 要求至少 f + 1 个节点已经持久化才能返回 commit 消息,且 * 选举新 leader 的基数至少是 f + 1 个节点; 则最多允许有 f 个 broker 节点崩溃(包括已经崩溃的旧 leader),此时剩余 f + 1,个,因为 f + 1 个节点已经持久化了消息,所以这两个集合至少有 1 个节点重合,所以从剩余的 f + 1 个节点中选举出的新 leader 必然包含所有消息,从而保证不丢消息。 >每条消息至少保存在 f + 1 个节点,一共 2f + 1 个节点,则从中任意选 f + 1 个节点,其中必然至少包含一个已经保存了消息的节点。

kun-song

# Kafka 可靠性配置 ## broker 配置 * `default.replication.factor` 默认值 1 * `min.insync.replicas` 默认值 1 * `unclean.leader.election.enable` 默认值 `false` ### `min.insync.replicas` 当 producer 配置 `acks=all/-1` 时,消息必须写入 `min.insync.replicas` 个副本才会被认为成功(committed),若 ISR 数量低于该配置,则生产者抛出 `NotEnoughReplicas` 或 `NotEnoughReplicasAfterAppend` 配置。 因此 `min.insync.replicas` 和 `acks=all` 联合使用时,可以提供更强的持久性保证(greater durability guarantees),典型配置如: * `replication.factor=3` * `min.insync.replicas=2` * `acks=all` 该配置下,若消息无法写入大部分副本(即 ISR 少于允许的最小值),则 producer 将抛出异常。 >问题:`min.insync.replicas=1` 不也可以吗? >答: >* 若为 1,则该 ISR 崩溃,且禁止 unclean election,则该分区不可用,直到(手动恢复)原 leader; >* 若为 2,则 leader 崩溃后,另外一个 ISR 将被选为新 leader,此时: > + consumer 可正常读取 > + producer 写入失败,但 producer 可以通过重试,直到 ISR 恢复 2 个,从而保证可用性; ### `unclean.leader.election.enable` 若所有 in-sync replicas 一直正常工作当然好,但最坏的情况总会发生:所有 in-sync replicas 停止工作,此时有两个选择: * 选择一致性,抛弃可用性:该分区停止接受写入请求,直到原来的 ISR 恢复,从而产生 leaer; * 选择可用性,抛弃一致性:将 out-of-sync replicas 选举为新 leader; 第二种选择可能导致消息丢失,因为新 leader 原本是 out-of-sync replica,不保证持有所有消息。 该配置在不同 Kafka 版本中默认值不同,但从 0.11.0.0 开始,默认 `false`,即禁用 unclean election, The new default favors durability over availability. Users who wish to to retain the previous behavior should set the broker config unclean.leader.election.enable to true. 2. ## producer 配置 * `retries` * `retry.backoff.ms` enable.idempotence ``` The amount of time to wait before attempting to retry a failed request to a given topic partition. This avoids repeatedly sending requests in a tight loop under some failure scenarios. 默认 100ms ```

kun-song

# Kafka 与 Quorum Kafka 使用的是 Quorum 的一种变体: * 用动态维护的 ISR 替代每次选举的 xx quorum

kun-song

https://github.com/confluentinc/kafka-streams-examples/issues/141 https://stackoverflow.com/questions/52438161/set-timestamp-in-output-with-kafka-streams

kun-song

https://stackoverflow.com/questions/46591456/recordtoolargeexception-in-kafka-streams-join http://bigdatadecode.club/kafka%20channel%E5%86%99%E5%85%A5kafka%E6%8A%A5RecordTooLargeException%E5%BC%82%E5%B8%B8.html https://community.hortonworks.com/articles/74001/kafka-producer-running-into-multiple-orgapachekafk.html

kun-song

when creating internal topics https://cwiki.apache.org/confluence/display/KAFKA/KIP-173%3A+Add+prefix+to+StreamsConfig+to+enable+setting+default+internal+topic+configs

kun-song

## Kafka Streams 配置 ### broker 端 消息最大值: * broker 级别:`message.max.bytes`,默认 1M; * topic 级别:`max.message.bytes`,默认 1M; 设置了消息最大值,还需要设置 replica 复制消息的最大值,否则 follow 可能永远无法追上 leader: * `replica.fetch.max.bytes=1048576` ### producer 端 producer 一个请求发送的最大字节数,必须小于 broker 配置的消息最大值: * `max.request.size` ### consumer 端 * `max.partition.fetch.bytes` ### streams 内部 topic streams 内部使用的 topic 的消息最大值设置: * `max.message.bytes` ```Java props.put(StreamsConfig.topicPrefix(TopicConfig.MAX_MESSAGE_BYTES_CONFIG), xx); ``` 创建内部 topic 时,将使用上面的配置限制消息的最大值。 streams 本身既是 consumer,又是 producer,所以还需要配置: * `max.partition.fetch.bytes` * `max.request.size` ```Java props.put(StreamsConfig.consumerPrefix(ConsumerConfig.MAX_PARTITION_FETCH_BYTES_CONFIG), xx); props.put(StreamsConfig.producerPrefix(ProducerConfig.MAX_REQUEST_SIZE_CONFIG), yy); ``` * `buffer.memory`:produer 缓存消息的内存大小,默认 33554432 bytes;

kun-song

# KIP-328: Ability to suppress updates for KTables https://cwiki.apache.org/confluence/display/KAFKA/KIP-328%3A+Ability+to+suppress+updates+for+KTables

kun-song

https://www.confluent.de/blog/kafka-streams-take-on-watermarks-and-triggers

kun-song

flush 页缓存落盘时间与丢消息:若 acks=1,且 leader 写入也缓存成功,但落盘之前挂掉,则丢消息。 heap 内存 vs 页缓存内存,堆设置越大,则页缓存能用的空间越小。

kun-song

老师好,在消息重试的时候,分区策略会重新再计算一次吗?比如一开始选择到5号分区,但是5号分区有问题导致重试,重试的时候可以重试发送到别的分区上吗? 作者回复: 不会的。消息重试只是简单地将消息重新发送到之前的分区 如果消费过程中出现rebalance,那么可能造成因果关系之消费了因后rebalance,然后不处理之前的partition了,后面的消费者也无法处理该partition的“果”,请问,您对这种情况怎么处理的呢? 作者回复: 可以试试sticky assignor,即设置consumer端参数partition.assignment.strategy=class org.apache.kafka.clients.consumer.StickyAssignor

kun-song

Kafka 丢失消息举例