Learn to Match: Two-Sided Matching with Temporally Extended Feedback

Please check the full paper here.

Github code