./one_shot.sh I cmn common_param: common_params_print_info: build 1184 (732dd501) with GNU 16.2.0 for Linux x86_64 I cmn common_param: common_params_print_info: verbosity = 4 (adjust with the `-lv N` CLI arg) I cmn common_param: device_info: I cmn common_param: - ROCm0 : AMD Radeon RX 6800 (16368 MiB, 16342 MiB free) I cmn common_param: - ROCm1 : AMD Radeon RX 6700 XT (12272 MiB, 12248 MiB free) I cmn common_param: - Vulkan0 : AMD Radeon RX 6800 (RADV NAVI21) (16368 MiB, 16252 MiB free) I cmn common_param: - Vulkan1 : AMD Radeon RX 6700 XT (RADV NAVI22) (12272 MiB, 12256 MiB free) I cmn common_param: - CPU : 11th Gen Intel(R) Core(TM) i5-11400F @ 2.60GHz (31965 MiB, 31965 MiB free) I cmn common_param: system_info: n_threads = 6 (n_threads_batch = 6) / 12 | ROCm : NO_VMM = 1 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | I srv init: running without SSL I srv init: using 11 threads for HTTP server W srv llama_server: ----------------- W srv llama_server: CORS is set to allow all origins ('*') and no API key is set W srv llama_server: this can be a security risk (cross-origin attacks) W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 W srv llama_server: ----------------- I srv start: binding port with default address family I srv load_model: loading model '/home/eaman/lm/models/gianni/qwen3.6-27b-IQ4_XS-pure-with-MTP-IQ4.gguf' I srv load_model: local path '/home/eaman/lm/models/gianni/qwen3.6-27b-IQ4_XS-pure-with-MTP-IQ4.gguf' I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15894 + (19782 = 13244 + 6443 + 94) + -19308 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15742 + (19934 = 13244 + 6592 + 96) + -19308 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I srv load_model: [spec] recurrent checkpoint fit reservation: device ROCm0, 149.62 MiB (context 6443.25 -> 6592.88 MiB) I srv load_model: [spec] recurrent checkpoint fit reservation: 149.62 MiB total I common_params_fit_impl: getting device memory data for initial parameters: I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (19782 = 13244 + 6443 + 94) + -19304 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I common_params_fit_impl: projected to use 19782 MiB of device memory vs. 15890 MiB of free device memory I common_params_fit_impl: cannot meet free memory target of 209 MiB, need to reduce device memory by 4102 MiB I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (13674 = 13244 + 395 + 34) + -13196 | I common_memory_breakdown_print: | - Host | 650 = 644 + 0 + 6 | I common_params_fit_impl: context size reduced from 262144 to 88832 -> need 4102 MiB less memory in total I common_params_fit_impl: entire model can be fit by reducing context I common_fit_params: successfully fit params to free device memory I common_fit_params: fitting params to free memory took 0.63 seconds I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 16190 + (13381 = 13244 + 97 + 38) + -13203 | I common_memory_breakdown_print: | - Host | 671 = 644 + 0 + 26 | I srv load_model: [spec] estimated memory usage of MTP context is 136.55 MiB I common_params_fit_impl: getting device memory data for initial parameters: I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (19782 = 13244 + 6443 + 94) + -19304 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I common_params_fit_impl: projected to use 19782 MiB of device memory vs. 15890 MiB of free device memory I common_params_fit_impl: cannot meet free memory target of 346 MiB, need to reduce device memory by 4238 MiB I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (13674 = 13244 + 395 + 34) + -13196 | I common_memory_breakdown_print: | - Host | 650 = 644 + 0 + 6 | I common_params_fit_impl: context size reduced from 262144 to 82944 -> need 4241 MiB less memory in total I common_params_fit_impl: entire model can be fit by reducing context I common_fit_params: successfully fit params to free device memory I common_fit_params: fitting params to free memory took 0.63 seconds I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 16190 + (13373 = 13244 + 91 + 37) + -13195 | I common_memory_breakdown_print: | - Host | 669 = 644 + 0 + 25 | I srv load_model: [spec] refined MTP memory estimate at n_ctx=82944 is 128.64 MiB I common_params_fit_impl: getting device memory data for initial parameters: I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (19782 = 13244 + 6443 + 94) + -19304 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I common_params_fit_impl: projected to use 19782 MiB of device memory vs. 15890 MiB of free device memory I common_params_fit_impl: cannot meet free memory target of 338 MiB, need to reduce device memory by 4231 MiB I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (13674 = 13244 + 395 + 34) + -13196 | I common_memory_breakdown_print: | - Host | 650 = 644 + 0 + 6 | I common_params_fit_impl: context size reduced from 262144 to 83200 -> need 4235 MiB less memory in total I common_params_fit_impl: entire model can be fit by reducing context I common_fit_params: successfully fit params to free device memory I common_fit_params: fitting params to free memory took 0.63 seconds I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 16190 + (13373 = 13244 + 91 + 37) + -13195 | I common_memory_breakdown_print: | - Host | 669 = 644 + 0 + 25 | I srv load_model: [spec] refined MTP memory estimate at n_ctx=83200 is 128.99 MiB I cmn common_init_: fitting params to device memory ... I cmn common_init_: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on) I common_params_fit_impl: getting device memory data for initial parameters: I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (19782 = 13244 + 6443 + 94) + -19304 | I common_memory_breakdown_print: | - Host | 713 = 644 + 0 + 69 | I common_params_fit_impl: projected to use 19782 MiB of device memory vs. 15890 MiB of free device memory I common_params_fit_impl: cannot meet free memory target of 338 MiB, need to reduce device memory by 4231 MiB I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | I common_memory_breakdown_print: | - ROCm0 (RX 6800) | 16368 = 15890 + (13674 = 13244 + 395 + 34) + -13196 | I common_memory_breakdown_print: | - Host | 650 = 644 + 0 + 6 | I common_params_fit_impl: context size reduced from 262144 to 83200 -> need 4235 MiB less memory in total I common_params_fit_impl: entire model can be fit by reducing context I common_fit_params: successfully fit params to free device memory I common_fit_params: fitting params to free memory took 0.63 seconds I llama_model_loader: loaded meta data with 52 key-value pairs and 866 tensors from /home/eaman/lm/models/gianni/qwen3.6-27b-IQ4_XS-pure-with-MTP-IQ4.gguf (version GGUF V3 (latest)) I llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. I llama_model_loader: - kv 0: general.architecture str = qwen35 I llama_model_loader: - kv 1: general.type str = model I llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 I llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000 I llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000 I llama_model_loader: - kv 5: general.name str = Qwen3.6-27B I llama_model_loader: - kv 6: general.basename str = Qwen3.6-27B I llama_model_loader: - kv 7: general.quantized_by str = Unsloth I llama_model_loader: - kv 8: general.size_label str = 27B I llama_model_loader: - kv 9: general.license str = apache-2.0 I llama_model_loader: - kv 10: general.license.link str = https://huggingface.co/Qwen/Qwen3.6-2... I llama_model_loader: - kv 11: general.repo_url str = https://huggingface.co/unsloth I llama_model_loader: - kv 12: general.base_model.count u32 = 1 I llama_model_loader: - kv 13: general.base_model.0.name str = Qwen3.6 27B I llama_model_loader: - kv 14: general.base_model.0.organization str = Qwen I llama_model_loader: - kv 15: general.base_model.0.repo_url str = https://huggingface.co/Qwen/Qwen3.6-27B I llama_model_loader: - kv 16: general.tags arr[str,2] = ["unsloth", "image-text-to-text"] I llama_model_loader: - kv 17: qwen35.block_count u32 = 65 I llama_model_loader: - kv 18: qwen35.context_length u32 = 262144 I llama_model_loader: - kv 19: qwen35.embedding_length u32 = 5120 I llama_model_loader: - kv 20: qwen35.feed_forward_length u32 = 17408 I llama_model_loader: - kv 21: qwen35.attention.head_count u32 = 24 I llama_model_loader: - kv 22: qwen35.attention.head_count_kv u32 = 4 I llama_model_loader: - kv 23: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0] I llama_model_loader: - kv 24: qwen35.rope.freq_base f32 = 10000000.000000 I llama_model_loader: - kv 25: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001 I llama_model_loader: - kv 26: qwen35.attention.key_length u32 = 256 I llama_model_loader: - kv 27: qwen35.attention.value_length u32 = 256 I llama_model_loader: - kv 28: qwen35.ssm.conv_kernel u32 = 4 I llama_model_loader: - kv 29: qwen35.ssm.state_size u32 = 128 I llama_model_loader: - kv 30: qwen35.ssm.group_count u32 = 16 I llama_model_loader: - kv 31: qwen35.ssm.time_step_rank u32 = 48 I llama_model_loader: - kv 32: qwen35.ssm.inner_size u32 = 6144 I llama_model_loader: - kv 33: qwen35.full_attention_interval u32 = 4 I llama_model_loader: - kv 34: qwen35.rope.dimension_count u32 = 64 I llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2 I llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35 I llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ... I llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... I llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",... I llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046 I llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248055 I llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044 I llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false I llama_model_loader: - kv 44: general.quantization_version u32 = 2 I llama_model_loader: - kv 45: general.file_type u32 = 30 I llama_model_loader: - kv 46: quantize.imatrix.file str = imatrix_unsloth.gguf_file I llama_model_loader: - kv 47: quantize.imatrix.dataset str = unsloth_calibration_Qwen3.6-27B.txt I llama_model_loader: - kv 48: quantize.imatrix.entries_count u32 = 496 I llama_model_loader: - kv 49: quantize.imatrix.chunks_count u32 = 76 I llama_model_loader: - kv 50: tokenizer.chat_template str = {# =========================\n Qwen ... I llama_model_loader: - kv 51: qwen35.nextn_predict_layers u32 = 1 I llama_model_loader: - type f32: 360 tensors I llama_model_loader: - type q8_0: 1 tensors I llama_model_loader: - type q4_K: 6 tensors I llama_model_loader: - type q5_K: 1 tensors I llama_model_loader: - type iq4_xs: 498 tensors I print_info: file format = GGUF V3 (latest) I print_info: file type = IQ4_XS - 4.25 bpw I print_info: file size = 13.56 GiB (4.26 BPW) I llama_prepare_model_devices: using device ROCm0 (AMD Radeon RX 6800) (0000:03:00.0) - 16190 MiB free I load: 0 unused tokens I load: printing all EOG tokens: I load: - 248044 ('<|endoftext|>') I load: - 248046 ('<|im_end|>') I load: - 248063 ('<|fim_pad|>') I load: - 248064 ('<|repo_name|>') I load: - 248065 ('<|file_sep|>') I load: special tokens cache size = 33 I load: token to piece cache size = 1.7581 MB I print_info: arch = qwen35 I print_info: vocab_only = 0 I print_info: no_alloc = 0 I print_info: n_ctx_train = 262144 I print_info: n_embd_inp = 5120 I print_info: n_embd = 5120 I print_info: n_embd_out = 5120 I print_info: n_layer = 64 I print_info: n_layer_all = 65 I print_info: n_head = 24 I print_info: n_head_kv = 4 I print_info: n_rot = 64 I print_info: n_swa = 0 I print_info: is_swa_any = 0 I print_info: n_embd_head_k = 256 I print_info: n_embd_head_v = 256 I print_info: n_gqa = 6 I print_info: n_embd_k_gqa = 1024 I print_info: n_embd_v_gqa = 1024 I print_info: f_norm_eps = 0.0e+00 I print_info: f_norm_rms_eps = 1.0e-06 I print_info: f_clamp_kqv = 0.0e+00 I print_info: f_max_alibi_bias = 0.0e+00 I print_info: f_logit_scale = 0.0e+00 I print_info: f_attn_scale = 0.0e+00 I print_info: f_attn_value_scale = 0.0000 I print_info: n_ff = 17408 I print_info: n_expert = 0 I print_info: n_expert_used = 0 I print_info: n_expert_groups = 0 I print_info: n_group_used = 0 I print_info: causal attn = 1 I print_info: pooling type = -1 I print_info: rope type = 40 I print_info: rope scaling = linear I print_info: freq_base_train = 10000000.0 I print_info: freq_scale_train = 1 I print_info: n_ctx_orig_yarn = 262144 I print_info: rope_yarn_log_mul = 0.0000 I print_info: rope_finetuned = unknown I print_info: mrope sections = [11, 11, 10, 0] I print_info: ssm_d_conv = 4 I print_info: ssm_d_inner = 6144 I print_info: ssm_d_state = 128 I print_info: ssm_dt_rank = 48 I print_info: ssm_n_group = 16 I print_info: ssm_dt_b_c_rms = 0 I print_info: model type = 27B I print_info: model params = 27.32 B I print_info: general.name = Qwen3.6-27B I print_info: vocab type = BPE I print_info: n_vocab = 248320 I print_info: n_merges = 247587 I print_info: BOS token = 248044 '<|endoftext|>' I print_info: EOS token = 248046 '<|im_end|>' I print_info: EOT token = 248046 '<|im_end|>' I print_info: PAD token = 248055 '<|vision_pad|>' I print_info: LF token = 198 'Ċ' I print_info: FIM PRE token = 248060 '<|fim_prefix|>' I print_info: FIM SUF token = 248062 '<|fim_suffix|>' I print_info: FIM MID token = 248061 '<|fim_middle|>' I print_info: FIM PAD token = 248063 '<|fim_pad|>' I print_info: FIM REP token = 248064 '<|repo_name|>' I print_info: FIM SEP token = 248065 '<|file_sep|>' I print_info: EOG token = 248044 '<|endoftext|>' I print_info: EOG token = 248046 '<|im_end|>' I print_info: EOG token = 248063 '<|fim_pad|>' I print_info: EOG token = 248064 '<|repo_name|>' I print_info: EOG token = 248065 '<|file_sep|>' I print_info: max token length = 256 I load_tensors: loading model tensors, this can take a while... (load_mode = none) I load_tensors: offloading output layer to GPU I load_tensors: offloading 64 repeating layers to GPU I load_tensors: offloaded 66/66 layers to GPU I load_tensors: ROCm0 model buffer size = 13244.73 MiB I load_tensors: ROCm_Host model buffer size = 644.14 MiB I cmn common_init_: added <|endoftext|> logit bias = -inf I cmn common_init_: added <|im_end|> logit bias = -inf I cmn common_init_: added <|fim_pad|> logit bias = -inf I cmn common_init_: added <|repo_name|> logit bias = -inf I cmn common_init_: added <|file_sep|> logit bias = -inf I llama_context: constructing llama_context I llama_context: n_seq_max = 1 I llama_context: n_ctx = 83200 I llama_context: n_ctx_seq = 83200 I llama_context: n_batch = 1024 I llama_context: n_ubatch = 128 I llama_context: causal_attn = 1 I llama_context: flash_attn = enabled I llama_context: pipeline mode = disabled I llama_context: kv_unified = false I llama_context: freq_base = 10000000.0 I llama_context: freq_scale = 1 I llama_context: n_rs_seq = 1 I llama_context: n_outputs_max = 33 I llama_context: n_outputs_max_per_seq = 33 I llama_context: n_ctx_seq (83200) < n_ctx_train (262144) -- the full capacity of the model will not be utilized I llama_context: ROCm_Host output buffer size = 0.95 MiB I llama_kv_cache: ROCm0 KV buffer size = 1950.00 MiB I llama_kv_cache: size = 1950.00 MiB ( 83200 cells, 16 layers, 1/1 seqs), K (q5_1): 975.00 MiB, V (q5_1): 975.00 MiB I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256 I llama_memory_recurrent: ROCm0 RS buffer size = 299.25 MiB I llama_memory_recurrent: size = 299.25 MiB ( 1 cells, 64 layers, 1 seqs 1 rs_seq), R (f32): 11.25 MiB, S (f32): 288.00 MiB I sched_reserve: reserving ... I resolve_fused_ops: resolving fused Gated Delta Net support: I resolve_fused_ops: fused Gated Delta Net (autoregressive) enabled I resolve_fused_ops: fused Gated Delta Net (chunked) enabled I resolve_fused_ops: resolving fused Lightning Indexer support: I resolve_fused_ops: Lightning Indexer enabled I resolve_fused_ops: resolving fused DeepSeek V4 HC support: I resolve_fused_ops: fused DeepSeek V4 HC pre enabled I resolve_fused_ops: fused DeepSeek V4 HC comb enabled I resolve_fused_ops: fused DeepSeek V4 HC post enabled I sched_reserve: ROCm0 compute buffer size = 51.08 MiB I sched_reserve: ROCm_Host compute buffer size = 25.58 MiB I sched_reserve: graph nodes = 3991 I sched_reserve: graph splits = 2 I sched_reserve: reserve took 17.13 ms, sched copies = 1 I cmn init: llama threadpool init, n_threads = 6 I common_speculative_init_result: creating MTP draft context against the target model '/home/eaman/lm/models/gianni/qwen3.6-27b-IQ4_XS-pure-with-MTP-IQ4.gguf' I llama_context: constructing llama_context I llama_context: n_seq_max = 1 I llama_context: n_ctx = 83200 I llama_context: n_ctx_seq = 83200 I llama_context: n_batch = 1024 I llama_context: n_ubatch = 128 I llama_context: causal_attn = 1 I llama_context: flash_attn = enabled I llama_context: pipeline mode = disabled I llama_context: kv_unified = false I llama_context: freq_base = 10000000.0 I llama_context: freq_scale = 1 I llama_context: n_rs_seq = 0 I llama_context: n_outputs_max = 1 I llama_context: n_outputs_max_per_seq = 1 I llama_context: n_ctx_seq (83200) < n_ctx_train (262144) -- the full capacity of the model will not be utilized I llama_context: ROCm_Host output buffer size = 0.95 MiB I llama_kv_cache: ROCm0 KV buffer size = 91.41 MiB I llama_kv_cache: size = 91.41 MiB ( 83200 cells, 1 layers, 1/1 seqs), K (q4_0): 45.70 MiB, V (q4_0): 45.70 MiB I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256 I sched_reserve: reserving ... I resolve_fused_ops: resolving fused Gated Delta Net support: I resolve_fused_ops: fused Gated Delta Net (autoregressive) enabled I resolve_fused_ops: fused Gated Delta Net (chunked) enabled I resolve_fused_ops: resolving fused Lightning Indexer support: I resolve_fused_ops: Lightning Indexer enabled I resolve_fused_ops: resolving fused DeepSeek V4 HC support: I resolve_fused_ops: fused DeepSeek V4 HC pre enabled I resolve_fused_ops: fused DeepSeek V4 HC comb enabled I resolve_fused_ops: fused DeepSeek V4 HC post enabled I sched_reserve: ROCm0 compute buffer size = 37.58 MiB I sched_reserve: ROCm_Host compute buffer size = 25.58 MiB I sched_reserve: graph nodes = 62 I sched_reserve: graph splits = 2 I sched_reserve: reserve took 8.94 ms, sched copies = 1 I cmn common_conte: the context supports bounded partial sequence removal I srv load_model: initializing, n_slots = 1, n_ctx_slot = 83200, kv_unified = 'false' I spec common_specu: adding speculative implementation 'ngram-mod' I spec common_specu: - n_match=24, n_max=32, n_min=8 I spec common_specu: - mod size=4194304 (16.000 MB) I spec common_specu: adding speculative implementation 'draft-mtp' I spec common_specu: - n_max=5, n_min=0, p_min=0.82, n_embd=5120, backend_sampling=1 I spec common_specu: - gpu_layers=-1, cache_k=q4_0, cache_v=q4_0, ctx_tgt=yes, ctx_dft=yes, devices=[default] W llama_sampler_backend_support: device 'ROCm0' does not have support for op TOP_K needed for sampler 'top-k' I sched_reserve: reserving ... I sched_reserve: ROCm0 compute buffer size = 37.58 MiB I sched_reserve: ROCm_Host compute buffer size = 25.58 MiB I sched_reserve: graph nodes = 64 I sched_reserve: graph splits = 2 I sched_reserve: reserve took 12.84 ms, sched copies = 1 I srv load_model: speculative decoding context initialized I slot load_model: id 0 | task -1 | new slot, n_ctx = 83200 I srv load_model: prompt cache is enabled, size limit: 6000 MiB I srv load_model: use `--cache-ram 0` to disable the prompt cache I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 I srv load_model: context checkpoints enabled, max = 128, min spacing = 8192 I srv init: idle slots will be saved to prompt cache upon starting a new task I srv init: init: chat template, example_format: '<|im_start|>system You are a helpful assistant<|im_end|> <|im_start|>user Hello<|im_end|> <|im_start|>assistant Hi there<|im_end|> <|im_start|>user How are you?<|im_end|> <|im_start|>assistant ' I srv init: init: chat template, thinking = 1 I srv llama_server: model loaded I srv llama_server: listening on http://0.0.0.0:8080 W srv llama_server: NOTICE: server default port will be changed to :9931 in a future release W srv llama_server: ref: https://github.com/ggml-org/llama.cpp/pull/26508 I srv update_slots: all slots are idle I srv server_strea: conv_id=i75raywxs6 (empty=0) I srv operator(): chat format: peg-native I slot get_availabl: id 0 | task -1 | - skipping, slot is empty I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 I srv get_availabl: updating prompt cache I srv load: - looking for better prompt, base f_keep = -1.000, f_sim = 0.000 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 6000.000 MiB, 83200 tokens, 6291456000 est) I srv get_availabl: prompt cache update took 0.01 ms I cmn common_reaso: activated, budget=6096 tokens I slot launch_slot_: id 0 | task -1 | sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> ?min-p -> ?xtc -> temp-ext -> dist I slot launch_slot_: id 0 | task -1 | sampler params: repeat_last_n = 64, repeat_penalty = 1.000, frequency_penalty = 0.000, presence_penalty = 0.000 dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = 64 top_k = 20, top_p = 0.950, min_p = 0.000, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.500 mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 I slot operator(): id 0 | task 0 | new prompt, n_ctx_slot = 83200, n_keep = 0, task.n_tokens = 531 I slot operator(): id 0 | task 0 | cached n_tokens = 0, memory_seq_rm [0, end) I slot operator(): id 0 | task 0 | cached n_tokens = 399, memory_seq_rm [399, end) I slot create_check: id 0 | task 0 | created context checkpoint 1 of 128 (pos_min = 398, pos_max = 398, n_tokens = 399, size = 150.072 MiB) I slot operator(): id 0 | task 0 | cached n_tokens = 486, memory_seq_rm [486, end) I slot create_check: id 0 | task 0 | created context checkpoint 2 of 128 (pos_min = 485, pos_max = 485, n_tokens = 486, size = 150.169 MiB) I slot operator(): id 0 | task 0 | cached n_tokens = 527, memory_seq_rm [527, end) I slot init_sampler: id 0 | task 0 | init sampler, took 0.06 ms, tokens: text = 531, total = 531 I slot create_check: id 0 | task 0 | created context checkpoint 3 of 128 (pos_min = 526, pos_max = 526, n_tokens = 527, size = 150.215 MiB) I spec begin: ngram_mod occupancy = 507/4194304 (0.00) I ~llama_io_write_device: allocated 'ROCm0' buffer 149.625 MiB I slot print_timing: id 0 | task 0 | n_gen = 110, tg = 36.19 t/s, tg_3s = 36.53 t/s I cmn common_reaso: deactivated (natural end) I slot print_timing: id 0 | task 0 | n_gen = 245, tg = 39.97 t/s, tg_3s = 43.64 t/s I slot print_timing: id 0 | task 0 | n_gen = 403, tg = 43.99 t/s, tg_3s = 52.08 t/s I slot print_timing: id 0 | task 0 | n_gen = 561, tg = 46.01 t/s, tg_3s = 52.08 t/s I slot print_timing: id 0 | task 0 | n_gen = 730, tg = 47.78 t/s, tg_3s = 54.79 t/s I slot print_timing: id 0 | task 0 | n_gen = 896, tg = 48.86 t/s, tg_3s = 54.24 t/s I slot print_timing: id 0 | task 0 | n_gen = 1163, tg = 53.78 t/s, tg_3s = 81.13 t/s I slot print_timing: id 0 | task 0 | n_gen = 1398, tg = 56.55 t/s, tg_3s = 75.89 t/s I slot print_timing: id 0 | task 0 | n_gen = 1538, tg = 55.34 t/s, tg_3s = 45.65 t/s I slot print_timing: id 0 | task 0 | n_gen = 1710, tg = 55.42 t/s, tg_3s = 56.16 t/s I slot print_timing: id 0 | task 0 | n_gen = 1963, tg = 57.92 t/s, tg_3s = 83.27 t/s I slot print_timing: id 0 | task 0 | n_gen = 2198, tg = 59.39 t/s, tg_3s = 75.27 t/s I slot print_timing: id 0 | task 0 | prompt eval time = 1816.86 ms / 531 tokens ( 3.42 ms per token, 292.26 tokens per second) I slot print_timing: id 0 | task 0 | eval time = 39605.41 ms / 2352 tokens ( 16.85 ms per token, 59.36 tokens per second) I slot print_timing: id 0 | task 0 | total time = 41422.28 ms / 2883 tokens I slot print_timing: id 0 | task 0 | graphs reused = 162 I slot print_timing: id 0 | task 0 | draft acceptance = 0.93741 ( 1857 accepted / 1981 generated), mean len = 5.73 I slot print_timing: id 0 | task 0 | acc per pos = (0.987, 0.804, 0.690, 0.565, 0.481, 0.053, 0.053, 0.053, 0.053, 0.053, 0.051, 0.051, 0.051, 0.048, 0.048, 0.048, 0.048, 0.043, 0.043, 0.043, 0.043, 0.043, 0.041, 0.041, 0.041, 0.041, 0.038, 0.038, 0.036, 0.036, 0.031, 0.028) I slot print_timing: id 0 | task 0 | MTP replays = 17 events / 242 tokens I spec common_specu: statistics ngram-mod: #calls(b,g,a) = 1 478 21, #gen drafts = 21, #acc drafts = 21, #gen tokens = 672, #acc tokens = 585, #mean acc len = 28.86, #acc rate/pos = (1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 1.000, 0.952, 0.952, 0.952, 0.905, 0.905, 0.905, 0.905, 0.810, 0.810, 0.810, 0.810, 0.810, 0.762, 0.762, 0.762, 0.762, 0.714, 0.714, 0.667, 0.667, 0.524), dur(b,g,a) = 0.027, 0.576, 0.007 ms I spec common_specu: statistics draft-mtp: #calls(b,g,a) = 1 457 372, #gen drafts = 372, #acc drafts = 369, #gen tokens = 1309, #acc tokens = 1289, #mean acc len = 4.47, #acc rate/pos = (0.992, 0.796, 0.675, 0.551, 0.452), dur(b,g,a) = 0.002, 4628.712, 0.425 ms I slot release: id 0 | task 0 | stop processing: n_tokens = 2883, truncated = 0 I srv update_slots: all slots are idle