Deep q learning

No.10949124 ViewReplyOriginalReport
I have been struggling for few weeks with params of a simulation where a biped try to learn how to walk. Especially on rewards.

Now this autistic biped found a way to jump as fast as he can on the floor... so exact opposite to what I want
I guess my rewards are fucked and make it do so but whatever.

Is the fact that the biped suicide every single time as fast as he can proving that this simulation can work if I change the reward policy (if we assume that the rewards used force him to do so) ?

Pic related