Item-level confidence is calculated by mapping individual tokens to parsed entity strings and computing both the average probability and minimum probability across that sequence, contrasting with raw token-level confidence which measures only single-token likelihood.