Understanding that Arabic has emphatic consonants and pharyngeal sounds (see Arabic Pronunciation) is the easy part. Actually producing and distinguishing them takes deliberate, specific practice.
A "minimal pair" is two words that differ by exactly one sound — like a plain and an emphatic consonant. Practicing minimal pairs back-to-back (listening, then repeating) trains your ear and mouth on the exact distinction that matters, far more efficiently than practicing random unrelated vocabulary and hoping the distinction sinks in eventually.
When first practicing a pharyngeal sound like ع or ح, deliberately over-exaggerate the throat constriction — it will sound unnatural at first, but exaggeration helps you feel where the sound is actually produced. Dialing it back to a natural level is much easier once you've found the right location than trying to approximate it timidly from the start.
Recording your own attempt at a word immediately after native audio, then comparing them side by side, reveals gaps your ear alone won't catch in the moment — self-perception of your own accent is notoriously unreliable without this kind of direct comparison.
Since a doubled consonant is a real length distinction, practice holding it noticeably longer than you'd expect — most learners under-exaggerate doubled consonants at first, since English doesn't have an equivalent contrast to calibrate against.
See also the reference page: Arabic Pronunciation.
Practicing minimal pairs — words that differ by exactly one sound, like a plain and emphatic consonant — trains the specific distinction directly, more efficiently than general vocabulary practice.
Deliberately exaggerate the throat constriction at first, even though it will sound unnatural — this helps you locate where the sound is produced, and it's easier to dial back to a natural level afterward than to approximate it timidly from the start.
Self-perception of your own pronunciation is notoriously unreliable in the moment — recording yourself and comparing directly against native audio reveals gaps you wouldn't otherwise notice.
Under-pronounce — since English has no equivalent length contrast to calibrate against, most learners need to deliberately exaggerate the extra length at first.