GGUF header of the DSpark draft file (file name DeepSeek-V4-Flash-0731-DSpark.gguf, 10,896,057,568 bytes), read 2026-09-26 with a plain-Python reader over the key-value section only (5,248,323 bytes read; no tensor data touched). magic b'GGUF', version 3, tensors 81, key-value pairs 62. Keys whose value is a long array or a string of tokenizer data are summarised. general.architecture = dflash tokenizer.chat_template = {%- if not add_generation_prompt is defined -%} {%- set add_generation_prompt = false -%} {%- endif -%} {%- if not ... general.type = model general.sampling.top_p = 1.0 general.sampling.temp = 1.0 general.name = DeepSeek V4 Flash 0731 general.version = 0731 general.basename = DeepSeek-V4-Flash general.size_label = 256x594M general.license = mit dflash.block_count = 3 dflash.context_length = 1048576 dflash.embedding_length = 4096 dflash.attention.head_count = 64 dflash.attention.head_count_kv = 1 dflash.rope.scaling.type = yarn dflash.rope.scaling.factor = 16.0 dflash.rope.scaling.original_context_length = 65536 dflash.rope.scaling.yarn_beta_fast = 32.0 dflash.rope.scaling.yarn_beta_slow = 1.0 dflash.rope.freq_base = 10000.0 dflash.attention.layer_norm_rms_epsilon = 9.999999974752427e-07 dflash.expert_count = 256 dflash.expert_used_count = 6 dflash.expert_gating_func = 4 dflash.attention.key_length = 512 dflash.attention.value_length = 512 general.file_type = 38 dflash.rope.dimension_count = 64 dflash.attention.q_lora_rank = 1024 dflash.attention.sliding_window = 128 dflash.expert_feed_forward_length = 2048 dflash.expert_shared_count = 1 dflash.expert_weights_scale = 1.5 dflash.expert_weights_norm = True dflash.swiglu_clamp_exp = [10.0, 10.0, 10.0] dflash.swiglu_clamp_shexp = [10.0, 10.0, 10.0] dflash.attention.indexer.head_count = 64 dflash.attention.indexer.key_length = 128 dflash.attention.indexer.top_k = 512 dflash.attention.output_group_count = 8 dflash.attention.output_lora_rank = 1024 dflash.attention.compress_ratios = [0, 0, 0] dflash.attention.compress_rope_freq_base = 160000.0 dflash.hyper_connection.count = 4 dflash.hyper_connection.sinkhorn_iterations = 20 dflash.hyper_connection.epsilon = 9.999999974752427e-07 dflash.hash_layer_count = 0 dflash.block_size = 5 dflash.target_layers = [41, 42, 43] general.quantization_version = 2 tokenizer.ggml.model = gpt2 tokenizer.ggml.pre = joyai-llm tokenizer.ggml.tokens = [array of 129280 values] tokenizer.ggml.token_type = [array of 129280 values] tokenizer.ggml.merges = [array of 127741 values] tokenizer.ggml.bos_token_id = 0 tokenizer.ggml.eos_token_id = 1 tokenizer.ggml.padding_token_id = 1 tokenizer.ggml.add_bos_token = False tokenizer.ggml.add_eos_token = False tokenizer.ggml.mask_token_id = 128799 has dflash.attention.sliding_window_pattern: False