179 questions
No questions match those filters.
A colleague wants to replace your PRF expansion with an...
This is one of the questions in the full AI/ML interview bank. Pro unlocks all 1789 questions; Premium includes the same bank plus the highest daily Practice limit.
See plansThe proposal fails on two critical fronts: the lack of a reward function and the inability to keep the policy synchronized with the rapidly changing index. Without click logs, there is no objective for RL training unless you mine weak pairs or purchase judgments. Furthermore, with hourly index churn, the reward function is effectively a moving target, meaning the model would always be trained against a stale corpus. The defensible path is to maintain the current PRF system while gathering data to eventually support a more advanced model.