• ExtraJudgement@lemmy.world
    link
    fedilink
    arrow-up
    1
    ·
    1 day ago

    If you provide storage to upload four audio .wav files (each about 400 KiB), I can send you samples generated by me speaking, then cloned with w-okada using a generic voice I believe is freely usable. It’s far from perfect, but it’s good enough for me to clearly tell which is which just by listening.

    • FishFace@piefed.social
      link
      fedilink
      English
      arrow-up
      1
      ·
      13 hours ago

      I don’t know of anywhere suitable, sorry. Also if you are not comfortable sharing your own voice (which I perfectly understand, I also would not want to) I don’t think a voice changer is the way to go.

      Maybe we will have to leave that thread here then - I also understand if you don’t want to go digging for examples. I would be interested to know still if you hear the glottal stops you’re talking about in the example above, though?

      Generalising the Spiegelei example: in order to read a word in a language which writes its compounds without spaces, you need to be able to break words down into their compounds correctly. Just reading the letters, with no knowledge of vocabulary, “Gewinnspiel” could be analysed in at least two ways: “Gewinns·piel” or “Gewinn·spiel”. The “s” would be pronounced differently between the two analyses (respectively as a /z/ or a /ʃ/ following regular German pronunciation rules).

      Fun stuff!

      • ExtraJudgement@lemmy.world
        link
        fedilink
        arrow-up
        1
        ·
        8 hours ago

        I would be interested to know still if you hear the glottal stops you’re talking about in the example above, though?

        In the voice changed output for “Das Kind”? It’s not perfect, but yes. If you’re interested in the “biomechanics”: It’s the difference between a miniscule stopping of the continuous exhalation from your mouth for the s before starting a new burst exhalation for the k vs doing both as part of the same single exhalation that directly changes from continuous to burst.

          • ExtraJudgement@lemmy.world
            link
            fedilink
            arrow-up
            1
            ·
            edit-2
            4 hours ago

            Yes, there’s a stop of exhalation between the s and the K. If there wasn’t, it would sound more like the sk of the English word sky instead.

            EDIT: He does not, however, stop between Kindern and werden and not between werden and Leute, which does sound weird to me. Less like standard German and more like a dialect.

            • FishFace@piefed.social
              link
              fedilink
              English
              arrow-up
              1
              ·
              3 hours ago

              Well, there has to be a stop of exhalation, because the “k” sound is a stop and that is basically the definition of a stop! One thing I want to be pedantic about though: this is not a glottal stop, which is a specific sound made by closing the glottis. “k” denotes a velar stop, which is different.

              You can see this stop as a dark area on the spectrogram:

              image

              But here’s the spectrogram of someone saying sky:

              image

              You can see there is the same dark area! Pronouncing a “k” requires cutting off the airstream in the middle of a word also. The duration is about the same; 30 or 40ms in both cases.

              If you look further through “aus Kindern werden Leute” in an audio editor you won’t see this dark area between “Kindern” and “werden”, nor between “werden” and “Leute”. These don’t have stops between them, so the sound is continuous, even though they’re separate words.

              It’s easy to see the stop in “Leute” - this is almost as complete as that between “aus” and “Kindern”. Less obvious are those in “Kindern” and “werden” (both for a letter “d”), but you can also see the dark areas.

              I hope this is enough to explain how what you’re hearing between some words also occurs within words, and between many words this phenomenon doesn’t happen. Phonetics is tricky. It’s like learning to draw - at first you don’t hear (or see) just what is there, but instead what you know is there. You know that “aus Kindern” or “das Kind” is two separate words, so you hear a break. You know that “Leute” or “sky” is one word, so you don’t hear any break.

              • ExtraJudgement@lemmy.world
                link
                fedilink
                arrow-up
                1
                ·
                2 hours ago

                I’ve corrected my previous comments w.r.t. the terminology, thank you.

                That being said, assuming those two spectrograms are using a comparable frequency scale, they show quite a bit of a difference: The upper spectrogram shows a near complete “black” stop across all but the lowest frequencies, while the lower spectrogram does not: The 2nd-lowest frequencies continue virtually unchanged, the mid frequencies are only reduced to a dark purple, not completely dark, and only the upper frequencies are truly dark. Those are not the same sound.

                • FishFace@piefed.social
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  1 hour ago

                  Well you’re right that there is colour in the “sky” spectrogram, but if you loop the dark segment (easy in audacity) you will hear that it’s just background noise. The “Kindern” sample has had noise removed or been recorded in a studio setting. It’s also worth bearing in mind the colours are on a log scale; if you switch over to looking at the waveform you can see just how little is going on in both of those little gaps.

                  If this interests you there’s a whole world of phonology and linguistics out there waiting for you… my mind was blown as a teenager when I learned about ich-Laut and ach-Laut.