In practice, professional transcription requires 5-10 times longer than the source material, depending on its content.
Slowing down to half speed and listening to it once isn't enough. You need a first pass to write all of it down which is usually not done with simple playback at half speed (which isn't that easy to understand) but instead with foot pedal switch for pausing and rewinding. It takes much longer than the video (well, depends on the video - different speakers speak at very different speeds), and at least a full pass re-listening everything for proofreading.
Technical videos take extra time - you often need to take minute or two to clarify a single term that you don't know, verify that you're not confusing it with another word and that it's spelled properly; and during an hour-long video such terms and the required time add up The same goes up for surnames - it takes a second to blurt out "paper by Mumblemumble et al", and it takes much longer to transcribe that even if the paper can be looked up in other related documents (and not always it can). A single neccessary clarification + a few related emails to solve it already can taka half an hour.
Captioning is some extra work in getting sure that the written segments align with the speech - it basically means that you have to note the start/end information, usually it's done at the initial transcript stage by the play/pause switches (the software packages apply the previously played segment start/end timestamps) and then you need to adjust many of them in proofreading. Not really an extra stage, but takes more work and care.
Transcribing at double speed, as you propose, is not really practical. People speak at 130-180wpm. Decent typists usually type at 65-75wpm. At that rate you would barely manage to 'type what you think you hear right now', which is usable for some purposes not really a transcript.