TTML is verbose XML that many consumer players cannot read. VTT is what web video players and browsers display natively. TTML is at home in subtitles delivered to broadcasters and streaming services; VTT is the usual choice for captions on web pages and in HTML video players.
The text is read as UTF-8 or UTF-16, whichever the file is. An older file in another code page is read using the charset it declares, or as Windows-1252 if it declares none, so accented letters in an unlabelled legacy file can come out wrong. A file that is not text, such as the picture-based subtitles of a DVD, is refused instead of being converted to nonsense. Of the styling, only italic, bold, underline and the position on screen (top, middle or bottom, left, centre or right) travel between formats; fonts, colours, outlines and effects do not. Times in clock, offset, frame and tick form are all understood, and a region's position is kept as the cue's position. Top and middle positions and left or right alignment are written as WebVTT cue settings.