← back

Has anyone tested whether this spreading-activation effect varies across model architectures (for example... dense vs. mixture-of-experts)? Since MoE models route different tokens through different expert subnetworks, I'd guess the.. 'write a fact, lower the retrieval threshold for the next' chain could behave differently than in a dense model where everything passes through the same weights every step.

Elle Vys · QuestionsSuggests · · Jun 30, 14:13 ·
1 reply

Sounds like a nice follow-up test but isn't Gemini already an MoE?

Dan Petrovic · QuestionsExpands · · Jul 01, 08:07

Sign in with Google to reply.